# EuroEval, OrangeSum (résumé)

> Fiche du benchmark EuroEval, OrangeSum (résumé) sur Quelle IA, édition du 28 septembre 2026. Le modèle résume un article de presse en français. En tête au 28 septembre 2026 : GLM-5.3-Flash (40,9 %). Notation, limites et classement des 69 modèles mesurés.

Page : https://quelleia.com/benchmarks/euroeval_orange_sum/
Source du benchmark : https://euroeval.com/leaderboards/Monolingual/french/
Mainteneur : EuroEval
Tâche : Français
État : actif

## À quoi il sert

Sur Quelle IA, EuroEval, OrangeSum (résumé) sert à noter la tâche Français : c'est l'un de ses 10 benchmarks.

## Comment c'est noté

La proximité entre son résumé et un résumé de référence, mesurée automatiquement.

## Ce qu'il ne dit pas

Un bon résumé différent de la référence est mal noté. Les écarts entre modèles sont faibles.

## Qui le tient

EuroEval.

## Comment lire le score

Chaque trait est un modèle, placé à son meilleur score. Le test reste très dur : d'après sa courbe d'étalonnage, même un modèle de score IA 100 n'y obtient qu'environ 39 %. Le meilleur, GLM-5.3-Flash, atteint 40,9 %. Le benchmark sera dit saturé quand un modèle atteindra 95 %.

Ce que le benchmark attend à chaque niveau, d'après sa courbe d'étalonnage :

- Score IA 51 (niveau de GPT-4o (Nov 2024)) : 37 %
- Score IA 76 (niveau de Claude Sonnet 4.5) : 38 %
- Score IA 100 (niveau de Claude Opus 5.5) : 39 %

## Son poids dans le score IA

0,62 % du score général. EuroEval, OrangeSum (résumé) porte 6 % de la tâche Français (10 % du score général), que se partagent 10 benchmarks.

## Classement (69 modèles mesurés · 91 mesures, édition du 28 septembre 2026)

Un modèle par ligne, à sa meilleure configuration. « Suggère » : le score IA que cette seule mesure indique.

| Rang | Modèle | Éditeur | Réflexion | Score publié | Suggère | Source |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GLM-5.3-Flash](https://quelleia.com/modeles/glm-5-3-flash.md) | Z.ai (Zhipu AI) | non précisée | 40,9 % | 143,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 2 | [Gemini 3 Pro](https://quelleia.com/modeles/gemini-3-pro.md) | Google DeepMind | non précisée | 39,5 % | 111,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 3 | [Claude Sonnet 4.5](https://quelleia.com/modeles/claude-sonnet-4-5.md) | Anthropic | Sans | 39,5 % | 111,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 4 | [Gemini 3.7 Flash](https://quelleia.com/modeles/gemini-3-7-flash.md) | Google DeepMind | non précisée | 39,2 % | 104,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 5 | [Apertus v1.5 70B](https://quelleia.com/modeles/apertus-v1-5-70b.md) | Swiss AI (EPFL, ETH Zurich) | non précisée | 39,1 % | 102,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 6 | [Gemma 4 31B IT](https://quelleia.com/modeles/gemma-4-31b-it.md) | Google DeepMind | non précisée | 39 % | 99,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 7 | [Gemma 4 26B A4B](https://quelleia.com/modeles/gemma-4-26b-a4b.md) | Google DeepMind | non précisée | 38,9 % | 97,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 8 | [GPT-4.1](https://quelleia.com/modeles/gpt-4-1.md) | OpenAI | non précisée | 38,9 % | 96,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 9 | [Gemini 3.5 Flash-Lite](https://quelleia.com/modeles/gemini-3-5-flash-lite.md) | Google DeepMind | non précisée | 38,8 % | 95,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 10 | [Gemini 3.1 Flash-Lite](https://quelleia.com/modeles/gemini-3-1-flash-lite.md) | Google DeepMind | non précisée | 38,6 % | 90,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 11 | [Gemini 3 Flash](https://quelleia.com/modeles/gemini-3-flash.md) | Google DeepMind | Sans | 38,6 % | 90,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 11 | [Gemini 2.5 Flash-Lite (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-flash-lite-jun-2025.md) | Google DeepMind | Sans | 38,6 % | 90,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 11 | [EuroLLM 9B Instruct (2512)](https://quelleia.com/modeles/eurollm-9b-instruct-2512.md) | EuroLLM | non précisée | 38,6 % | 90,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 14 | [Mistral Small 3.1](https://quelleia.com/modeles/mistral-small-3-1.md) | Mistral AI | non précisée | 38,6 % | 88,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 15 | [EuroLLM 9B Instruct](https://quelleia.com/modeles/eurollm-9b-instruct.md) | EuroLLM | non précisée | 38,5 % | 86,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 16 | [Llama 3.3 70B](https://quelleia.com/modeles/llama-3-3-70b.md) | Meta AI | non précisée | 38,4 % | 84,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 17 | [Claude Haiku 4.5](https://quelleia.com/modeles/claude-haiku-4-5.md) | Anthropic | non précisée | 38,4 % | 84,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 17 | [Luciole 23B Instruct 1.1](https://quelleia.com/modeles/luciole-23b-instruct-1-1.md) | OpenLLM-France | non précisée | 38,4 % | 84,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 19 | [Qwen 3.8 27B](https://quelleia.com/modeles/qwen-3-8-27b.md) | Alibaba | non précisée | 38,3 % | 82,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 20 | [Grok 4.1 Fast](https://quelleia.com/modeles/grok-4-1-fast.md) | xAI | Élevée | 38,2 % | 80,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 21 | [Gemini 3.6 Flash](https://quelleia.com/modeles/gemini-3-6-flash.md) | Google DeepMind | non précisée | 38,2 % | 80,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 22 | [GPT-5.4 Mini](https://quelleia.com/modeles/gpt-5-4-mini.md) | OpenAI | Sans | 38 % | 74,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 22 | [Qwen3.6 27B](https://quelleia.com/modeles/qwen3-6-27b.md) | Alibaba | non précisée | 38 % | 74,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 24 | [Ministral 3 14B](https://quelleia.com/modeles/ministral-3-14b.md) | Mistral AI | non précisée | 38 % | 74,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 24 | [Qwen2.5-72B](https://quelleia.com/modeles/qwen2-5-72b.md) | Alibaba | non précisée | 38 % | 74,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 26 | [Qwen3-32B](https://quelleia.com/modeles/qwen3-32b.md) | Alibaba | Sans | 38 % | 74,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 27 | [Llama 3.1-8B](https://quelleia.com/modeles/llama-3-1-8b.md) | Meta AI | non précisée | 37,9 % | 73,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 28 | [Gemma 3 27B](https://quelleia.com/modeles/gemma-3-27b.md) | Google DeepMind | non précisée | 37,9 % | 72,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 28 | [Luciole 8B Instruct 1.1](https://quelleia.com/modeles/luciole-8b-instruct-1-1.md) | OpenLLM-France | non précisée | 37,9 % | 72,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 30 | [Qwen3-14B](https://quelleia.com/modeles/qwen3-14b.md) | Alibaba | Sans | 37,9 % | 71,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 30 | [Gemma 3 12B](https://quelleia.com/modeles/gemma-3-12b.md) | Google DeepMind | non précisée | 37,9 % | 71,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 32 | [Mistral Small 3.2](https://quelleia.com/modeles/mistral-small-3-2.md) | Mistral AI | non précisée | 37,9 % | 71,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 33 | [Qwen3.5-2B](https://quelleia.com/modeles/qwen3-5-2b.md) | Alibaba | non précisée | 37,8 % | 71,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 34 | [granite-4.0-micro](https://quelleia.com/modeles/granite-4-0-micro.md) | IBM | non précisée | 37,8 % | 69,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 35 | [Claude Sonnet 4.6](https://quelleia.com/modeles/claude-sonnet-4-6.md) | Anthropic | non précisée | 37,8 % | 69,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 36 | [GPT-5 mini](https://quelleia.com/modeles/gpt-5-mini.md) | OpenAI | non précisée | 37,5 % | 63,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 37 | [Qwen3-8B](https://quelleia.com/modeles/qwen3-8b.md) | Alibaba | Sans | 37,5 % | 63,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 38 | [Gemma 3n E4B](https://quelleia.com/modeles/gemma-3n-e4b.md) | Google DeepMind | non précisée | 37,4 % | 60,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 39 | [GPT-5.6 Terra](https://quelleia.com/modeles/gpt-5-6-terra.md) | OpenAI | non précisée | 37,4 % | 60,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 39 | [Qwen3.5-35B-A3B](https://quelleia.com/modeles/qwen3-5-35b-a3b.md) | Alibaba | non précisée | 37,4 % | 60,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 41 | [Qwen3-4B](https://quelleia.com/modeles/qwen3-4b.md) | Alibaba | Sans | 37,4 % | 60,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 42 | [GPT-5.6 Luna](https://quelleia.com/modeles/gpt-5-6-luna.md) | OpenAI | non précisée | 37,4 % | 59,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 43 | [GPT-5.6 Sol](https://quelleia.com/modeles/gpt-5-6-sol.md) | OpenAI | non précisée | 37,3 % | 58,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 44 | [GPT-6 Astra](https://quelleia.com/modeles/gpt-6-astra.md) | OpenAI | non précisée | 37,3 % | 57,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 45 | [Ministral 3 8B](https://quelleia.com/modeles/ministral-3-8b.md) | Mistral AI | non précisée | 37,2 % | 55,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 46 | [GPT-5.4](https://quelleia.com/modeles/gpt-5-4.md) | OpenAI | Sans | 37,1 % | 53,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 47 | [Llama 3.2 3B](https://quelleia.com/modeles/llama-3-2-3b.md) | Meta AI | non précisée | 37,1 % | 52,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 48 | [Qwen3.5-9B](https://quelleia.com/modeles/qwen3-5-9b.md) | Alibaba | non précisée | 37 % | 51,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 49 | [Gemma 3 4B](https://quelleia.com/modeles/gemma-3-4b.md) | Google DeepMind | non précisée | 36,9 % | 48,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 50 | [GPT-5.2](https://quelleia.com/modeles/gpt-5-2.md) | OpenAI | non précisée | 36,8 % | 45,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 51 | [Qwen3.5-0.8B](https://quelleia.com/modeles/qwen3-5-0-8b.md) | Alibaba | non précisée | 36,5 % | 39,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 52 | [Qwen3-1.7B](https://quelleia.com/modeles/qwen3-1-7b.md) | Alibaba | Sans | 36,5 % | 38,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 53 | [Voxtral Small](https://quelleia.com/modeles/voxtral-small.md) | Mistral AI | non précisée | 36,4 % | 37,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 54 | [glm-4.7-flash](https://quelleia.com/modeles/glm-4-7-flash.md) | Z.ai (Zhipu AI) | non précisée | 36,3 % | 34,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 55 | [OLMo 3.1 32B Instruct](https://quelleia.com/modeles/olmo-3-1-32b-instruct.md) | Allen Institute for AI | non précisée | 36,2 % | 31,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 56 | [EuroLLM 22B Instruct](https://quelleia.com/modeles/eurollm-22b-instruct.md) | EuroLLM | non précisée | 36 % | 26,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 57 | [GPT-5.4 Nano](https://quelleia.com/modeles/gpt-5-4-nano.md) | OpenAI | Sans | 35,7 % | 19,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 58 | [Qwen3.5-4B](https://quelleia.com/modeles/qwen3-5-4b.md) | Alibaba | non précisée | 35,7 % | 17,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 59 | [Ministral 3 3B](https://quelleia.com/modeles/ministral-3-3b.md) | Mistral AI | Élevée | 35,1 % | 4,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 60 | [GPT-5 nano](https://quelleia.com/modeles/gpt-5-nano.md) | OpenAI | non précisée | 35 % | 2,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 61 | [Qwen3-0.6B](https://quelleia.com/modeles/qwen3-0-6b.md) | Alibaba | Sans | 34,8 % | -2,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 62 | [Nemotron 3 Nano 30B](https://quelleia.com/modeles/nemotron-3-nano-30b.md) | NVIDIA | non précisée | 34,7 % | -6,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 63 | [gpt-oss-20b](https://quelleia.com/modeles/gpt-oss-20b.md) | OpenAI | Faible | 34,6 % | -8,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 64 | [gpt-oss-120b](https://quelleia.com/modeles/gpt-oss-120b.md) | OpenAI | non précisée | 34,6 % | -9,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 65 | [Gemma 3 270M](https://quelleia.com/modeles/gemma-3-270m.md) | Google DeepMind | non précisée | 32,3 % | -67,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 66 | [Llama 3.2 1B](https://quelleia.com/modeles/llama-3-2-1b.md) | Meta AI | non précisée | 31,4 % | -89,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 67 | [Gemma 3 1B](https://quelleia.com/modeles/gemma-3-1b.md) | Google DeepMind | non précisée | 31 % | -102,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 68 | [Llama 3.1-70B](https://quelleia.com/modeles/llama-3-1-70b.md) | Meta AI | non précisée | 30,6 % | -111,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 69 | [Salamandra 7B Instruct](https://quelleia.com/modeles/salamandra-7b-instruct.md) | Barcelona Supercomputing Center | non précisée | 28,8 % | -160,3 | https://euroeval.com/leaderboards/Monolingual/french/ |

## En bref

**Que mesure EuroEval, OrangeSum (résumé) ?** Le modèle résume un article de presse en français. Sur Quelle IA, il est rangé dans la tâche Français.

**Comment EuroEval, OrangeSum (résumé) est-il noté ?** La proximité entre son résumé et un résumé de référence, mesurée automatiquement. Le test reste très dur : d'après sa courbe d'étalonnage, même un modèle de score IA 100 n'y obtient qu'environ 39 %.

**Qui est en tête sur EuroEval, OrangeSum (résumé) ?** GLM-5.3-Flash (Z.ai (Zhipu AI)) est en tête d'EuroEval, OrangeSum (résumé) avec 40,9 %, dans l'édition du 28 septembre 2026. Suivent Gemini 3 Pro (39,5 %) et Claude Sonnet 4.5 (39,5 %). 69 modèles y sont mesurés.

**Quelles sont les limites d'EuroEval, OrangeSum (résumé) ?** Un bon résumé différent de la référence est mal noté. Les écarts entre modèles sont faibles. Au 28 septembre 2026, le benchmark est actif.

**Combien pèse EuroEval, OrangeSum (résumé) dans le score IA ?** EuroEval, OrangeSum (résumé) compte pour 0,62 % du score général, dans l'édition du 28 septembre 2026. Il porte 6 % de la tâche Français (10 % du score général), que se partagent 10 benchmarks.

**Qui maintient EuroEval, OrangeSum (résumé) ?** EuroEval. Sa page de référence : euroeval.com/leaderboards/Monolingual/french.

## Les autres benchmarks de la tâche Français

- [Arena, questions en français](https://quelleia.com/benchmarks/arena_francais.md) : 224 modèles · 2,5 % du score
- [compar:IA](https://quelleia.com/benchmarks/comparia.md) : 114 modèles · 2,5 % du score
- [EuroEval, ScaLA (grammaire)](https://quelleia.com/benchmarks/euroeval_scala_fr.md) : 136 modèles · 0,62 % du score
- [EuroEval, Allociné (sentiment)](https://quelleia.com/benchmarks/euroeval_allocine.md) : 134 modèles · 0,62 % du score, saturé
- [EuroEval, ELTeC (noms propres)](https://quelleia.com/benchmarks/euroeval_eltec.md) : 134 modèles · 0,62 % du score
- [EuroEval, FQuAD (compréhension)](https://quelleia.com/benchmarks/euroeval_fquad.md) : 134 modèles · 0,62 % du score
- [EuroEval, HellaSwag (bon sens)](https://quelleia.com/benchmarks/euroeval_hellaswag_fr.md) : 134 modèles · 0,62 % du score
- [EuroEval, INCLUDE (connaissances)](https://quelleia.com/benchmarks/euroeval_include_fr.md) : 64 modèles · 0,62 % du score
- [EuroEval, MultiLoKo (connaissances locales)](https://quelleia.com/benchmarks/euroeval_multiloko_fr.md) : 64 modèles · 0,62 % du score

Source : Quelle IA, édition du 28 septembre 2026. https://quelleia.com
