# EuroEval, HellaSwag (bon sens)

> Fiche du benchmark EuroEval, HellaSwag (bon sens) sur Quelle IA, édition du 28 septembre 2026. Le modèle choisit la suite la plus plausible d'une situation décrite en français. En tête au 28 septembre 2026 : GPT-6 Astra (90,7 %). Notation, limites et classement des 134 modèles mesurés.

Page : https://quelleia.com/benchmarks/euroeval_hellaswag_fr/
Source du benchmark : https://euroeval.com/leaderboards/Monolingual/french/
Mainteneur : EuroEval
Tâche : Français
État : actif

## À quoi il sert

Sur Quelle IA, EuroEval, HellaSwag (bon sens) sert à noter la tâche Français : c'est l'un de ses 10 benchmarks. Il départage surtout les modèles dont le score IA approche 38.

## Comment c'est noté

La part de bonnes réponses, corrigée du hasard.

## Ce qu'il ne dit pas

Épreuve traduite de l'anglais, proche de la saturation pour les meilleurs modèles.

## Qui le tient

EuroEval.

## Comment lire le score

Chaque trait est un modèle, placé à son meilleur score. Un modèle qui obtient 50 % a un score IA d'environ 38. Le meilleur, GPT-6 Astra, atteint 90,7 %. Le benchmark sera dit saturé quand un modèle atteindra 95 %.

Ce que le benchmark attend à chaque niveau, d'après sa courbe d'étalonnage :

- Score IA 51 (niveau de GPT-4o (Nov 2024)) : 60 %
- Score IA 76 (niveau de Claude Sonnet 4.5) : 76 %
- Score IA 100 (niveau de Claude Opus 5.5) : 87 %

## Son poids dans le score IA

0,62 % du score général. EuroEval, HellaSwag (bon sens) porte 6 % de la tâche Français (10 % du score général), que se partagent 10 benchmarks.

## Classement (134 modèles mesurés · 190 mesures, édition du 28 septembre 2026)

Un modèle par ligne, à sa meilleure configuration. « Suggère » : le score IA que cette seule mesure indique.

| Rang | Modèle | Éditeur | Réflexion | Score publié | Suggère | Source |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-6 Astra](https://quelleia.com/modeles/gpt-6-astra.md) | OpenAI | non précisée | 90,7 % | 112,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 2 | [Gemini 3.7 Flash](https://quelleia.com/modeles/gemini-3-7-flash.md) | Google DeepMind | non précisée | 90,2 % | 110,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 3 | [Claude Sonnet 4.5](https://quelleia.com/modeles/claude-sonnet-4-5.md) | Anthropic | Sans | 87,9 % | 102,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 4 | [Qwen3-235B-A22B-Instruct (Jul 2025)](https://quelleia.com/modeles/qwen3-235b-a22b-instruct-jul-2025.md) | Alibaba | non précisée | 85,7 % | 96,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 5 | [Qwen3.6 27B](https://quelleia.com/modeles/qwen3-6-27b.md) | Alibaba | non précisée | 84,5 % | 93,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 6 | [Claude Sonnet 4.6](https://quelleia.com/modeles/claude-sonnet-4-6.md) | Anthropic | non précisée | 83,7 % | 91,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 7 | [Mistral Small 3](https://quelleia.com/modeles/mistral-small-3.md) | Mistral AI | non précisée | 83,4 % | 90,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 8 | [Mistral Small 3.1](https://quelleia.com/modeles/mistral-small-3-1.md) | Mistral AI | non précisée | 83,4 % | 90,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 9 | [Gemini 3.6 Flash](https://quelleia.com/modeles/gemini-3-6-flash.md) | Google DeepMind | non précisée | 82,6 % | 88,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 10 | [Qwen3-235B-A22B-Thinking (Jul 2025)](https://quelleia.com/modeles/qwen3-235b-a22b-thinking-jul-2025.md) | Alibaba | non précisée | 82,4 % | 88,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 11 | [Gemini 3 Pro](https://quelleia.com/modeles/gemini-3-pro.md) | Google DeepMind | non précisée | 81,8 % | 87,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 12 | [Gemini 3 Flash](https://quelleia.com/modeles/gemini-3-flash.md) | Google DeepMind | Sans | 81,1 % | 85,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 13 | [GLM-5.3-Flash](https://quelleia.com/modeles/glm-5-3-flash.md) | Z.ai (Zhipu AI) | non précisée | 81,1 % | 85,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 14 | [GPT-5.6 Sol](https://quelleia.com/modeles/gpt-5-6-sol.md) | OpenAI | non précisée | 80,9 % | 85,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 15 | [chutes/Qwen3-Next-80B-A3B-Instruct](https://quelleia.com/modeles/chutes-qwen3-next-80b-a3b-instruct.md) | Alibaba | non précisée | 80,4 % | 84,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 16 | [Llama 3.1-405B](https://quelleia.com/modeles/llama-3-1-405b.md) | Meta AI | non précisée | 80 % | 83,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 17 | [Gemini 2.5 Pro (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-pro-jun-2025.md) | Google DeepMind | non précisée | 79,9 % | 83,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 18 | [Gemma 4 31B IT](https://quelleia.com/modeles/gemma-4-31b-it.md) | Google DeepMind | non précisée | 79,5 % | 82,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 19 | [Qwen 3.8 27B](https://quelleia.com/modeles/qwen-3-8-27b.md) | Alibaba | non précisée | 79,4 % | 82,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 20 | [Mistral Small 3.2](https://quelleia.com/modeles/mistral-small-3-2.md) | Mistral AI | non précisée | 78,5 % | 80,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 21 | [GPT-5.4](https://quelleia.com/modeles/gpt-5-4.md) | OpenAI | Sans | 78 % | 79,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 22 | [GPT-5.6 Terra](https://quelleia.com/modeles/gpt-5-6-terra.md) | OpenAI | non précisée | 77,9 % | 79,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 23 | [Qwen2.5-72B](https://quelleia.com/modeles/qwen2-5-72b.md) | Alibaba | non précisée | 77,8 % | 79,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 24 | [Qwen3-Next-80B-A3B Thinking](https://quelleia.com/modeles/qwen3-next-80b-a3b-thinking.md) | Alibaba | Élevée | 77,4 % | 78,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 25 | [Voxtral Small](https://quelleia.com/modeles/voxtral-small.md) | Mistral AI | non précisée | 77 % | 77,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 26 | [GPT-5](https://quelleia.com/modeles/gpt-5.md) | OpenAI | Élevée | 76,4 % | 76,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 27 | [Magistral Small 1.2](https://quelleia.com/modeles/magistral-small-1-2.md) | Mistral AI | non précisée | 76,4 % | 76,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 28 | [Qwen3.5-35B-A3B](https://quelleia.com/modeles/qwen3-5-35b-a3b.md) | Alibaba | non précisée | 76 % | 75,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 29 | [Llama 4 Scout](https://quelleia.com/modeles/llama-4-scout.md) | Meta AI | non précisée | 75,9 % | 75,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 30 | [Llama 3.3 70B](https://quelleia.com/modeles/llama-3-3-70b.md) | Meta AI | non précisée | 75,7 % | 75,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 31 | [Yi-1.5-34B](https://quelleia.com/modeles/yi-1-5-34b.md) | 01.AI | non précisée | 75,3 % | 74,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 32 | [Llama-3.1-Nemotron-70B-Instruct](https://quelleia.com/modeles/llama-3-1-nemotron-70b-instruct.md) | NVIDIA | non précisée | 74,6 % | 73,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 33 | [Qwen3-32B](https://quelleia.com/modeles/qwen3-32b.md) | Alibaba | Sans | 74,4 % | 72,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 34 | [Aya Expanse 32B](https://quelleia.com/modeles/aya-expanse-32b.md) | Cohere | non précisée | 74 % | 72,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 35 | [Ministral 3 14B](https://quelleia.com/modeles/ministral-3-14b.md) | Mistral AI | Élevée | 73,9 % | 72,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 36 | [Qwen3-235B-A22B](https://quelleia.com/modeles/qwen3-235b-a22b.md) | Alibaba | non précisée | 73,7 % | 71,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 37 | [Gemini 2.5 Flash (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-flash-jun-2025.md) | Google DeepMind | Sans | 73,2 % | 70,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 38 | [GPT-4.1](https://quelleia.com/modeles/gpt-4-1.md) | OpenAI | non précisée | 72,8 % | 70,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 39 | [Gemini 3.1 Flash-Lite](https://quelleia.com/modeles/gemini-3-1-flash-lite.md) | Google DeepMind | non précisée | 72,2 % | 69,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 40 | [GPT-5.6 Luna](https://quelleia.com/modeles/gpt-5-6-luna.md) | OpenAI | non précisée | 71,9 % | 68,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 41 | [GPT-5.2](https://quelleia.com/modeles/gpt-5-2.md) | OpenAI | non précisée | 71,5 % | 68,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 42 | [Gemini 3.5 Flash-Lite](https://quelleia.com/modeles/gemini-3-5-flash-lite.md) | Google DeepMind | non précisée | 71,5 % | 68,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 43 | [Claude 3.7 Sonnet](https://quelleia.com/modeles/claude-3-7-sonnet.md) | Anthropic | Élevée | 71,4 % | 67,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 44 | [GPT-5.4 Mini](https://quelleia.com/modeles/gpt-5-4-mini.md) | OpenAI | Moyenne | 70,7 % | 66,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 45 | [Llama 3.1-70B](https://quelleia.com/modeles/llama-3-1-70b.md) | Meta AI | non précisée | 70,6 % | 66,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 46 | [Gemma 4 26B A4B](https://quelleia.com/modeles/gemma-4-26b-a4b.md) | Google DeepMind | non précisée | 70,4 % | 66,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 47 | [Claude 3.5 Sonnet](https://quelleia.com/modeles/claude-3-5-sonnet.md) | Anthropic | non précisée | 70,3 % | 66,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 48 | [Gemma 3 27B](https://quelleia.com/modeles/gemma-3-27b.md) | Google DeepMind | non précisée | 69,6 % | 65,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 49 | [QwQ-32B](https://quelleia.com/modeles/qwq-32b.md) | Alibaba | non précisée | 69,5 % | 64,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 50 | [Grok 3](https://quelleia.com/modeles/grok-3.md) | xAI | non précisée | 69 % | 64,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 51 | [Claude Haiku 4.5](https://quelleia.com/modeles/claude-haiku-4-5.md) | Anthropic | non précisée | 68,7 % | 63,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 52 | [Qwen3.5-9B](https://quelleia.com/modeles/qwen3-5-9b.md) | Alibaba | non précisée | 68,6 % | 63,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 53 | [GPT-5 mini](https://quelleia.com/modeles/gpt-5-mini.md) | OpenAI | Élevée | 67,9 % | 62,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 54 | [Ministral 3 8B](https://quelleia.com/modeles/ministral-3-8b.md) | Mistral AI | Élevée | 67,8 % | 62,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 55 | [o3](https://quelleia.com/modeles/o3.md) | OpenAI | non précisée | 67,4 % | 61,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 56 | [Gemma 2 27B](https://quelleia.com/modeles/gemma-2-27b.md) | Google DeepMind | non précisée | 67,2 % | 61,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 57 | [Qwen3-14B](https://quelleia.com/modeles/qwen3-14b.md) | Alibaba | Sans | 66,7 % | 60,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 58 | [Grok-2 (Dec 2024)](https://quelleia.com/modeles/grok-2-dec-2024.md) | xAI | non précisée | 66,3 % | 60,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 59 | [Phi-4](https://quelleia.com/modeles/phi-4.md) | Microsoft | non précisée | 65,9 % | 59,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 60 | [GLM-4.5-Air](https://quelleia.com/modeles/glm-4-5-air.md) | Z.ai (Zhipu AI) | Sans | 65,4 % | 58,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 61 | [Qwen3-30B-A3B](https://quelleia.com/modeles/qwen3-30b-a3b.md) | Alibaba | Sans | 65,2 % | 58,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 62 | [Gemma 3 12B](https://quelleia.com/modeles/gemma-3-12b.md) | Google DeepMind | non précisée | 64,8 % | 58,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 63 | [Llama 3-70B](https://quelleia.com/modeles/llama-3-70b.md) | Meta AI | non précisée | 63,4 % | 55,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 64 | [Magistral Small 1.0](https://quelleia.com/modeles/magistral-small-1-0.md) | Mistral AI | non précisée | 63 % | 55,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 65 | [Ministral 8B](https://quelleia.com/modeles/ministral-8b.md) | Mistral AI | non précisée | 63 % | 55,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 66 | [o4-mini](https://quelleia.com/modeles/o4-mini.md) | OpenAI | non précisée | 62,6 % | 54,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 67 | [Grok 4.1 Fast](https://quelleia.com/modeles/grok-4-1-fast.md) | xAI | Élevée | 61,8 % | 53,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 68 | [Gemma 2 9B](https://quelleia.com/modeles/gemma-2-9b.md) | Google DeepMind | non précisée | 61,5 % | 53,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 69 | [gpt-oss-120b](https://quelleia.com/modeles/gpt-oss-120b.md) | OpenAI | non précisée | 60,7 % | 52,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 70 | [GPT-4.1 mini](https://quelleia.com/modeles/gpt-4-1-mini.md) | OpenAI | non précisée | 60,4 % | 51,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 71 | [Qwen3.5-4B](https://quelleia.com/modeles/qwen3-5-4b.md) | Alibaba | non précisée | 60,1 % | 51,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 72 | [GPT-5 nano](https://quelleia.com/modeles/gpt-5-nano.md) | OpenAI | Élevée | 59,8 % | 50,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 73 | [Gemini 2.5 Flash-Lite (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-flash-lite-jun-2025.md) | Google DeepMind | Sans | 58,9 % | 49,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 74 | [qwen3-4b-instruct-2507](https://quelleia.com/modeles/qwen3-4b-instruct-2507.md) | Alibaba | non précisée | 58,9 % | 49,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 75 | [Apertus v1.5 70B](https://quelleia.com/modeles/apertus-v1-5-70b.md) | Swiss AI (EPFL, ETH Zurich) | non précisée | 57,1 % | 47,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 76 | [Apertus 70B Instruct](https://quelleia.com/modeles/apertus-70b-instruct.md) | Swiss AI (EPFL, ETH Zurich) | non précisée | 56,3 % | 46,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 77 | [Nemotron 3 Nano 30B](https://quelleia.com/modeles/nemotron-3-nano-30b.md) | NVIDIA | Sans | 56,2 % | 46,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 78 | [Qwen3-8B](https://quelleia.com/modeles/qwen3-8b.md) | Alibaba | Sans | 56,1 % | 45,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 79 | [Ministral 3 3B](https://quelleia.com/modeles/ministral-3-3b.md) | Mistral AI | Élevée | 56 % | 45,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 80 | [Qwen3-4B](https://quelleia.com/modeles/qwen3-4b.md) | Alibaba | Élevée | 56 % | 45,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 81 | [gpt-oss-20b](https://quelleia.com/modeles/gpt-oss-20b.md) | OpenAI | Moyenne | 55,7 % | 45,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 82 | [Gemma 3n E4B](https://quelleia.com/modeles/gemma-3n-e4b.md) | Google DeepMind | non précisée | 55,7 % | 45,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 83 | [DeepSeek-R1-Distill-Qwen-32B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-32b.md) | DeepSeek | non précisée | 55,4 % | 45,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 84 | [DeepSeek-R1-Distill-Llama-70B](https://quelleia.com/modeles/deepseek-r1-distill-llama-70b.md) | DeepSeek | non précisée | 54,4 % | 43,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 85 | [Grok-3 mini](https://quelleia.com/modeles/grok-3-mini.md) | xAI | Élevée | 53,6 % | 42,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 86 | [glm-4.7-flash](https://quelleia.com/modeles/glm-4-7-flash.md) | Z.ai (Zhipu AI) | non précisée | 53,6 % | 42,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 87 | [o3-mini](https://quelleia.com/modeles/o3-mini.md) | OpenAI | non précisée | 53,3 % | 42,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 88 | [EuroLLM 22B Instruct](https://quelleia.com/modeles/eurollm-22b-instruct.md) | EuroLLM | non précisée | 52 % | 40,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 89 | [Qwen1.5-14B](https://quelleia.com/modeles/qwen1-5-14b.md) | Alibaba | non précisée | 51,1 % | 39,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 90 | [Reka Flash 3](https://quelleia.com/modeles/reka-flash-3.md) | Reka AI | non précisée | 49,9 % | 37,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 91 | [Mistral NeMo](https://quelleia.com/modeles/mistral-nemo.md) | Mistral AI | non précisée | 49,5 % | 37,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 92 | [Claude 3.5 Haiku](https://quelleia.com/modeles/claude-3-5-haiku.md) | Anthropic | non précisée | 49,2 % | 36,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 93 | [c4ai-command-r-08-2024](https://quelleia.com/modeles/c4ai-command-r-08-2024.md) | Cohere | non précisée | 48,9 % | 36,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 94 | [Mixtral 8x7B](https://quelleia.com/modeles/mixtral-8x7b.md) | Mistral AI | non précisée | 48,6 % | 36,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 95 | [DeepSeek-R1-Distill-Qwen-14B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-14b.md) | DeepSeek | non précisée | 47,2 % | 34,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 96 | [OLMo 3.1 32B Instruct](https://quelleia.com/modeles/olmo-3-1-32b-instruct.md) | Allen Institute for AI | non précisée | 47,2 % | 34,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 97 | [granite-4.0-micro](https://quelleia.com/modeles/granite-4-0-micro.md) | IBM | non précisée | 46,6 % | 33,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 98 | [GPT-5.4 Nano](https://quelleia.com/modeles/gpt-5-4-nano.md) | OpenAI | Moyenne | 45,2 % | 31,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 99 | [Llama 3.1-8B](https://quelleia.com/modeles/llama-3-1-8b.md) | Meta AI | non précisée | 44,7 % | 30,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 100 | [Luciole 23B Instruct 1.1](https://quelleia.com/modeles/luciole-23b-instruct-1-1.md) | OpenLLM-France | non précisée | 44,1 % | 30,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 101 | [EuroLLM 9B Instruct (2512)](https://quelleia.com/modeles/eurollm-9b-instruct-2512.md) | EuroLLM | non précisée | 43,4 % | 29,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 102 | [Llama 2-70B](https://quelleia.com/modeles/llama-2-70b.md) | Meta AI | non précisée | 42,1 % | 27,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 103 | [deepseek-r1-0528-qwen3-8b](https://quelleia.com/modeles/deepseek-r1-0528-qwen3-8b.md) | DeepSeek | non précisée | 40,7 % | 25,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 104 | [nvidia-nemotron-nano-9b-v2](https://quelleia.com/modeles/nvidia-nemotron-nano-9b-v2.md) | NVIDIA | non précisée | 40,3 % | 25,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 105 | [Llama 3-8B](https://quelleia.com/modeles/llama-3-8b.md) | Meta AI | non précisée | 38,6 % | 22,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 106 | [Apertus 8B Instruct](https://quelleia.com/modeles/apertus-8b-instruct.md) | Swiss AI (EPFL | non précisée | 38 % | 21,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 107 | [Qwen3.5-2B](https://quelleia.com/modeles/qwen3-5-2b.md) | Alibaba | non précisée | 37 % | 20,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 108 | [OLMo 2 Furious 13B](https://quelleia.com/modeles/olmo-2-furious-13b.md) | Allen Institute for AI | non précisée | 36,5 % | 19,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 109 | [Qwen2.5-1.5B](https://quelleia.com/modeles/qwen2-5-1-5b.md) | Alibaba | non précisée | 34,3 % | 16,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 110 | [Gemma 3 4B](https://quelleia.com/modeles/gemma-3-4b.md) | Google DeepMind | non précisée | 33,9 % | 16,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 111 | [Phi-4 Reasoning](https://quelleia.com/modeles/phi-4-reasoning.md) | Microsoft | Élevée | 30,8 % | 11,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 112 | [Qwen3-1.7B](https://quelleia.com/modeles/qwen3-1-7b.md) | Alibaba | Sans | 29,7 % | 9,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 113 | [EuroLLM 9B Instruct](https://quelleia.com/modeles/eurollm-9b-instruct.md) | EuroLLM | non précisée | 28,5 % | 7,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 114 | [Falcon 2 11B](https://quelleia.com/modeles/falcon-2-11b.md) | Technology Innovation Institute | non précisée | 27,8 % | 6,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 115 | [Luciole 8B Instruct 1.1](https://quelleia.com/modeles/luciole-8b-instruct-1-1.md) | OpenLLM-France | non précisée | 27,7 % | 6,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 116 | [ALIA-40b](https://quelleia.com/modeles/alia-40b.md) | Barcelona Supercomputing Center | non précisée | 26,6 % | 4,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 117 | [Gemma 7B](https://quelleia.com/modeles/gemma-7b.md) | Google DeepMind | non précisée | 26,4 % | 4,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 118 | [Mistral 7B v0.2](https://quelleia.com/modeles/mistral-7b-v0-2.md) | Mistral AI | non précisée | 24,3 % | 0,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 119 | [Llama 3.2 3B](https://quelleia.com/modeles/llama-3-2-3b.md) | Meta AI | non précisée | 23,7 % | -0,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 120 | [DeepSeek-R1-Distill-Llama-8B](https://quelleia.com/modeles/deepseek-r1-distill-llama-8b.md) | DeepSeek | non précisée | 19,9 % | -7,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 121 | [Phi-4 Mini](https://quelleia.com/modeles/phi-4-mini.md) | Microsoft | Élevée | 19,7 % | -8,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 122 | [Mistral 7B v0.3](https://quelleia.com/modeles/mistral-7b-v0-3.md) | Mistral AI | non précisée | 18,2 % | -11,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 123 | [DeepSeek-R1-Distill-Qwen-7B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-7b.md) | DeepSeek | non précisée | 17,5 % | -12,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 124 | [Mistral 7B v0.1](https://quelleia.com/modeles/mistral-7b-v0-1.md) | Mistral AI | non précisée | 16,6 % | -15,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 125 | [Llama 2-13B](https://quelleia.com/modeles/llama-2-13b.md) | Meta AI | non précisée | 11,4 % | -29,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 126 | [Qwen3-0.6B](https://quelleia.com/modeles/qwen3-0-6b.md) | Alibaba | non précisée | 11,1 % | -30,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 127 | [Qwen3.5-0.8B](https://quelleia.com/modeles/qwen3-5-0-8b.md) | Alibaba | non précisée | 7,6 % | -43,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 128 | [Salamandra 7B Instruct](https://quelleia.com/modeles/salamandra-7b-instruct.md) | Barcelona Supercomputing Center | non précisée | 7,4 % | -45,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 129 | [Gemma 3 1B](https://quelleia.com/modeles/gemma-3-1b.md) | Google DeepMind | non précisée | 6,8 % | -48,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 130 | [DeepSeek-R1-Distill-Qwen-1.5B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-1-5b.md) | DeepSeek | non précisée | 5,1 % | -57,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 131 | [Phi-2](https://quelleia.com/modeles/phi-2.md) | Microsoft | non précisée | 4,6 % | -61,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 132 | [Gemma 2B](https://quelleia.com/modeles/gemma-2b.md) | Google DeepMind | non précisée | 3 % | -76,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 133 | [Llama 3.2 1B](https://quelleia.com/modeles/llama-3-2-1b.md) | Meta AI | non précisée | 0,5 % | -112,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 134 | [Gemma 3 270M](https://quelleia.com/modeles/gemma-3-270m.md) | Google DeepMind | non précisée | -0,3 % | -112,8 | https://euroeval.com/leaderboards/Monolingual/french/ |

## En bref

**Que mesure EuroEval, HellaSwag (bon sens) ?** Le modèle choisit la suite la plus plausible d'une situation décrite en français. Sur Quelle IA, il est rangé dans la tâche Français.

**Comment EuroEval, HellaSwag (bon sens) est-il noté ?** La part de bonnes réponses, corrigée du hasard. Un modèle qui obtient 50 % a un score IA d'environ 38.

**Qui est en tête sur EuroEval, HellaSwag (bon sens) ?** GPT-6 Astra (OpenAI) est en tête d'EuroEval, HellaSwag (bon sens) avec 90,7 %, dans l'édition du 28 septembre 2026. Suivent Gemini 3.7 Flash (90,2 %) et Claude Sonnet 4.5 (87,9 %). 134 modèles y sont mesurés.

**Quelles sont les limites d'EuroEval, HellaSwag (bon sens) ?** Épreuve traduite de l'anglais, proche de la saturation pour les meilleurs modèles. Au 28 septembre 2026, le benchmark est actif.

**Combien pèse EuroEval, HellaSwag (bon sens) dans le score IA ?** EuroEval, HellaSwag (bon sens) compte pour 0,62 % du score général, dans l'édition du 28 septembre 2026. Il porte 6 % de la tâche Français (10 % du score général), que se partagent 10 benchmarks.

**Qui maintient EuroEval, HellaSwag (bon sens) ?** EuroEval. Sa page de référence : euroeval.com/leaderboards/Monolingual/french.

## Les autres benchmarks de la tâche Français

- [Arena, questions en français](https://quelleia.com/benchmarks/arena_francais.md) : 224 modèles · 2,5 % du score
- [compar:IA](https://quelleia.com/benchmarks/comparia.md) : 114 modèles · 2,5 % du score
- [EuroEval, ScaLA (grammaire)](https://quelleia.com/benchmarks/euroeval_scala_fr.md) : 136 modèles · 0,62 % du score
- [EuroEval, Allociné (sentiment)](https://quelleia.com/benchmarks/euroeval_allocine.md) : 134 modèles · 0,62 % du score, saturé
- [EuroEval, ELTeC (noms propres)](https://quelleia.com/benchmarks/euroeval_eltec.md) : 134 modèles · 0,62 % du score
- [EuroEval, FQuAD (compréhension)](https://quelleia.com/benchmarks/euroeval_fquad.md) : 134 modèles · 0,62 % du score
- [EuroEval, OrangeSum (résumé)](https://quelleia.com/benchmarks/euroeval_orange_sum.md) : 69 modèles · 0,62 % du score
- [EuroEval, INCLUDE (connaissances)](https://quelleia.com/benchmarks/euroeval_include_fr.md) : 64 modèles · 0,62 % du score
- [EuroEval, MultiLoKo (connaissances locales)](https://quelleia.com/benchmarks/euroeval_multiloko_fr.md) : 64 modèles · 0,62 % du score

Source : Quelle IA, édition du 28 septembre 2026. https://quelleia.com
