# EuroEval, ScaLA (grammaire)

> Fiche du benchmark EuroEval, ScaLA (grammaire) sur Quelle IA, édition du 28 septembre 2026. Le modèle doit dire si une phrase française est correcte ou si elle a été abîmée. En tête au 28 septembre 2026 : GLM-5.3-Flash (78 %). Notation, limites et classement des 136 modèles mesurés.

Page : https://quelleia.com/benchmarks/euroeval_scala_fr/
Source du benchmark : https://euroeval.com/leaderboards/Monolingual/french/
Mainteneur : EuroEval
Tâche : Français
État : actif

## À quoi il sert

Sur Quelle IA, EuroEval, ScaLA (grammaire) sert à noter la tâche Français : c'est l'un de ses 10 benchmarks. Il départage surtout les modèles dont le score IA approche 67.

## Comment c'est noté

Un score d'accord avec la bonne réponse, de 0 à 100, corrigé du hasard.

## Ce qu'il ne dit pas

Juge la détection des fautes. La qualité de ce que le modèle écrit lui-même n'est pas notée.

## Qui le tient

EuroEval.

## Comment lire le score

Chaque trait est un modèle, placé à son meilleur score. Un modèle qui obtient 50 % a un score IA d'environ 67. Le meilleur, GLM-5.3-Flash, atteint 78 %. Le benchmark sera dit saturé quand un modèle atteindra 95 %.

Ce que le benchmark attend à chaque niveau, d'après sa courbe d'étalonnage :

- Score IA 51 (niveau de GPT-4o (Nov 2024)) : 43 %
- Score IA 76 (niveau de Claude Sonnet 4.5) : 55 %
- Score IA 100 (niveau de Claude Opus 5.5) : 65 %

## Son poids dans le score IA

0,62 % du score général. EuroEval, ScaLA (grammaire) porte 6 % de la tâche Français (10 % du score général), que se partagent 10 benchmarks.

## Classement (136 modèles mesurés · 192 mesures, édition du 28 septembre 2026)

Un modèle par ligne, à sa meilleure configuration. « Suggère » : le score IA que cette seule mesure indique.

| Rang | Modèle | Éditeur | Réflexion | Score publié | Suggère | Source |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GLM-5.3-Flash](https://quelleia.com/modeles/glm-5-3-flash.md) | Z.ai (Zhipu AI) | non précisée | 78 % | 133,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 2 | [Gemini 3 Pro](https://quelleia.com/modeles/gemini-3-pro.md) | Google DeepMind | non précisée | 75,6 % | 126,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 3 | [Gemma 4 31B IT](https://quelleia.com/modeles/gemma-4-31b-it.md) | Google DeepMind | non précisée | 74,2 % | 122,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 4 | [GPT-6 Astra](https://quelleia.com/modeles/gpt-6-astra.md) | OpenAI | non précisée | 74,2 % | 122,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 5 | [Ministral 3 14B](https://quelleia.com/modeles/ministral-3-14b.md) | Mistral AI | Élevée | 73,6 % | 121,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 6 | [Gemini 3.1 Flash-Lite](https://quelleia.com/modeles/gemini-3-1-flash-lite.md) | Google DeepMind | non précisée | 72,6 % | 118,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 7 | [Gemini 3 Flash](https://quelleia.com/modeles/gemini-3-flash.md) | Google DeepMind | Sans | 72,4 % | 117,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 8 | [Claude Sonnet 4.6](https://quelleia.com/modeles/claude-sonnet-4-6.md) | Anthropic | non précisée | 72,1 % | 117,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 9 | [Gemma 4 26B A4B](https://quelleia.com/modeles/gemma-4-26b-a4b.md) | Google DeepMind | non précisée | 70,8 % | 113,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 10 | [Gemini 3.7 Flash](https://quelleia.com/modeles/gemini-3-7-flash.md) | Google DeepMind | non précisée | 70,5 % | 112,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 11 | [Gemini 3.5 Flash-Lite](https://quelleia.com/modeles/gemini-3-5-flash-lite.md) | Google DeepMind | non précisée | 69,8 % | 111,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 12 | [Gemini 3.6 Flash](https://quelleia.com/modeles/gemini-3-6-flash.md) | Google DeepMind | non précisée | 69,3 % | 109,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 13 | [Apertus v1.5 70B](https://quelleia.com/modeles/apertus-v1-5-70b.md) | Swiss AI (EPFL, ETH Zurich) | non précisée | 69,1 % | 109,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 14 | [Qwen 3.8 27B](https://quelleia.com/modeles/qwen-3-8-27b.md) | Alibaba | non précisée | 68,7 % | 108,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 15 | [Ministral 3 8B](https://quelleia.com/modeles/ministral-3-8b.md) | Mistral AI | non précisée | 68,3 % | 107,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 16 | [GPT-5.2](https://quelleia.com/modeles/gpt-5-2.md) | OpenAI | non précisée | 67,5 % | 105,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 17 | [GPT-5.4 Mini](https://quelleia.com/modeles/gpt-5-4-mini.md) | OpenAI | Faible | 66,8 % | 103,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 18 | [Llama 3.1-70B](https://quelleia.com/modeles/llama-3-1-70b.md) | Meta AI | non précisée | 66,6 % | 103,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 19 | [GPT-5.4](https://quelleia.com/modeles/gpt-5-4.md) | OpenAI | Sans | 65 % | 99,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 20 | [Nemotron 3 Nano 30B](https://quelleia.com/modeles/nemotron-3-nano-30b.md) | NVIDIA | non précisée | 61,6 % | 91,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 21 | [GPT-5.6 Luna](https://quelleia.com/modeles/gpt-5-6-luna.md) | OpenAI | non précisée | 61,2 % | 90,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 22 | [Qwen3.5-35B-A3B](https://quelleia.com/modeles/qwen3-5-35b-a3b.md) | Alibaba | non précisée | 60,4 % | 89,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 23 | [Qwen3.5-4B](https://quelleia.com/modeles/qwen3-5-4b.md) | Alibaba | non précisée | 59,5 % | 87,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 24 | [Qwen3.5-9B](https://quelleia.com/modeles/qwen3-5-9b.md) | Alibaba | non précisée | 58,6 % | 85,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 25 | [glm-4.7-flash](https://quelleia.com/modeles/glm-4-7-flash.md) | Z.ai (Zhipu AI) | non précisée | 58 % | 84,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 26 | [Qwen3.6 27B](https://quelleia.com/modeles/qwen3-6-27b.md) | Alibaba | non précisée | 57,6 % | 83,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 27 | [Ministral 3 3B](https://quelleia.com/modeles/ministral-3-3b.md) | Mistral AI | non précisée | 56,7 % | 81,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 28 | [EuroLLM 22B Instruct](https://quelleia.com/modeles/eurollm-22b-instruct.md) | EuroLLM | non précisée | 56 % | 79,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 29 | [granite-4.0-micro](https://quelleia.com/modeles/granite-4-0-micro.md) | IBM | non précisée | 55,9 % | 79,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 30 | [Luciole 23B Instruct 1.1](https://quelleia.com/modeles/luciole-23b-instruct-1-1.md) | OpenLLM-France | non précisée | 55,1 % | 77,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 31 | [Gemma 3n E4B](https://quelleia.com/modeles/gemma-3n-e4b.md) | Google DeepMind | non précisée | 54 % | 75,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 32 | [EuroLLM 9B Instruct (2512)](https://quelleia.com/modeles/eurollm-9b-instruct-2512.md) | EuroLLM | non précisée | 53,6 % | 74,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 33 | [Llama 3.1-405B](https://quelleia.com/modeles/llama-3-1-405b.md) | Meta AI | non précisée | 52,6 % | 72,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 34 | [Qwen3.5-2B](https://quelleia.com/modeles/qwen3-5-2b.md) | Alibaba | non précisée | 52,3 % | 71,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 35 | [Gemma 3 4B](https://quelleia.com/modeles/gemma-3-4b.md) | Google DeepMind | non précisée | 50,6 % | 68,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 36 | [OLMo 3.1 32B Instruct](https://quelleia.com/modeles/olmo-3-1-32b-instruct.md) | Allen Institute for AI | non précisée | 50,4 % | 67,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 37 | [GPT-5.6 Terra](https://quelleia.com/modeles/gpt-5-6-terra.md) | OpenAI | non précisée | 50,2 % | 67,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 37 | [Grok 4.1 Fast](https://quelleia.com/modeles/grok-4-1-fast.md) | xAI | Élevée | 50,2 % | 67,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 39 | [Qwen3-1.7B](https://quelleia.com/modeles/qwen3-1-7b.md) | Alibaba | Sans | 49,2 % | 65,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 40 | [GPT-5.4 Nano](https://quelleia.com/modeles/gpt-5-4-nano.md) | OpenAI | Faible | 46,4 % | 59,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 41 | [Qwen2.5-72B](https://quelleia.com/modeles/qwen2-5-72b.md) | Alibaba | non précisée | 46,1 % | 58,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 42 | [Mistral Small 3.2](https://quelleia.com/modeles/mistral-small-3-2.md) | Mistral AI | non précisée | 45,9 % | 58,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 43 | [Gemini 2.5 Flash (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-flash-jun-2025.md) | Google DeepMind | Élevée | 45,6 % | 57,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 44 | [GPT-5.6 Sol](https://quelleia.com/modeles/gpt-5-6-sol.md) | OpenAI | non précisée | 45,4 % | 57,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 45 | [Qwen3-235B-A22B-Instruct (Jul 2025)](https://quelleia.com/modeles/qwen3-235b-a22b-instruct-jul-2025.md) | Alibaba | non précisée | 45,3 % | 57,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 46 | [Magistral Small 1.2](https://quelleia.com/modeles/magistral-small-1-2.md) | Mistral AI | non précisée | 45,1 % | 56,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 47 | [Claude Sonnet 4.5](https://quelleia.com/modeles/claude-sonnet-4-5.md) | Anthropic | Sans | 44,7 % | 55,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 48 | [chutes/Qwen3-Next-80B-A3B-Instruct](https://quelleia.com/modeles/chutes-qwen3-next-80b-a3b-instruct.md) | Alibaba | non précisée | 44,6 % | 55,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 49 | [Gemma 3 27B](https://quelleia.com/modeles/gemma-3-27b.md) | Google DeepMind | non précisée | 44,4 % | 55,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 50 | [Yi-1.5-34B](https://quelleia.com/modeles/yi-1-5-34b.md) | 01.AI | non précisée | 44,4 % | 55,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 51 | [Llama 3.1-8B](https://quelleia.com/modeles/llama-3-1-8b.md) | Meta AI | non précisée | 44,3 % | 54,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 52 | [Mixtral 8x7B](https://quelleia.com/modeles/mixtral-8x7b.md) | Mistral AI | non précisée | 43,2 % | 52,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 53 | [GPT-4.1](https://quelleia.com/modeles/gpt-4-1.md) | OpenAI | non précisée | 43 % | 51,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 54 | [Llama 3-70B](https://quelleia.com/modeles/llama-3-70b.md) | Meta AI | non précisée | 42,9 % | 51,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 55 | [Claude 3.7 Sonnet](https://quelleia.com/modeles/claude-3-7-sonnet.md) | Anthropic | Sans | 42,6 % | 51,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 56 | [Gemini 2.5 Pro (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-pro-jun-2025.md) | Google DeepMind | non précisée | 42,6 % | 51,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 57 | [GPT-5](https://quelleia.com/modeles/gpt-5.md) | OpenAI | Faible | 42,5 % | 50,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 58 | [Llama-3.1-Nemotron-70B-Instruct](https://quelleia.com/modeles/llama-3-1-nemotron-70b-instruct.md) | NVIDIA | non précisée | 42,2 % | 50,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 59 | [Qwen3-32B](https://quelleia.com/modeles/qwen3-32b.md) | Alibaba | Sans | 42,1 % | 50,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 60 | [gpt-oss-120b](https://quelleia.com/modeles/gpt-oss-120b.md) | OpenAI | non précisée | 42 % | 49,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 61 | [Gemini 2.5 Flash-Lite (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-flash-lite-jun-2025.md) | Google DeepMind | Élevée | 41,8 % | 49,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 61 | [Llama 4 Scout](https://quelleia.com/modeles/llama-4-scout.md) | Meta AI | non précisée | 41,8 % | 49,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 63 | [Gemma 2 27B](https://quelleia.com/modeles/gemma-2-27b.md) | Google DeepMind | non précisée | 41,7 % | 49,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 64 | [Voxtral Small](https://quelleia.com/modeles/voxtral-small.md) | Mistral AI | non précisée | 41,6 % | 49,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 65 | [Phi-4](https://quelleia.com/modeles/phi-4.md) | Microsoft | non précisée | 40,9 % | 47,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 66 | [Magistral Small 1.0](https://quelleia.com/modeles/magistral-small-1-0.md) | Mistral AI | non précisée | 40,9 % | 47,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 67 | [Qwen3-14B](https://quelleia.com/modeles/qwen3-14b.md) | Alibaba | Sans | 40,8 % | 47,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 68 | [c4ai-command-r-08-2024](https://quelleia.com/modeles/c4ai-command-r-08-2024.md) | Cohere | non précisée | 40,7 % | 47,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 69 | [Llama 3.3 70B](https://quelleia.com/modeles/llama-3-3-70b.md) | Meta AI | non précisée | 40,6 % | 46,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 70 | [GPT-4.1 mini](https://quelleia.com/modeles/gpt-4-1-mini.md) | OpenAI | non précisée | 40,6 % | 46,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 71 | [GLM-4.5-Air](https://quelleia.com/modeles/glm-4-5-air.md) | Z.ai (Zhipu AI) | Élevée | 40,2 % | 45,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 71 | [Mistral Small 3.1](https://quelleia.com/modeles/mistral-small-3-1.md) | Mistral AI | non précisée | 40,2 % | 45,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 73 | [Claude Haiku 4.5](https://quelleia.com/modeles/claude-haiku-4-5.md) | Anthropic | non précisée | 40,1 % | 45,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 74 | [Grok 3](https://quelleia.com/modeles/grok-3.md) | xAI | non précisée | 40 % | 45,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 75 | [Qwen3-30B-A3B](https://quelleia.com/modeles/qwen3-30b-a3b.md) | Alibaba | Sans | 39,9 % | 45,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 76 | [Gemma 3 12B](https://quelleia.com/modeles/gemma-3-12b.md) | Google DeepMind | non précisée | 39,7 % | 44,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 77 | [Gemma 2 9B](https://quelleia.com/modeles/gemma-2-9b.md) | Google DeepMind | non précisée | 39,5 % | 44,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 78 | [o3-mini](https://quelleia.com/modeles/o3-mini.md) | OpenAI | non précisée | 39,1 % | 43,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 79 | [Gemini 2.5 Flash (Apr 2025)](https://quelleia.com/modeles/gemini-2-5-flash-apr-2025.md) | Google DeepMind | non précisée | 39 % | 43,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 80 | [Gemini 1.5 Flash (Sep 2024)](https://quelleia.com/modeles/gemini-1-5-flash-sep-2024.md) | Google DeepMind | non précisée | 38,7 % | 42,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 81 | [Qwen1.5-14B](https://quelleia.com/modeles/qwen1-5-14b.md) | Alibaba | non précisée | 38,7 % | 42,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 82 | [Aya Expanse 32B](https://quelleia.com/modeles/aya-expanse-32b.md) | Cohere | non précisée | 38,2 % | 41,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 83 | [qwen3-4b-instruct-2507](https://quelleia.com/modeles/qwen3-4b-instruct-2507.md) | Alibaba | non précisée | 38,1 % | 41,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 84 | [Qwen3-8B](https://quelleia.com/modeles/qwen3-8b.md) | Alibaba | non précisée | 38 % | 41,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 85 | [Mistral NeMo](https://quelleia.com/modeles/mistral-nemo.md) | Mistral AI | non précisée | 37,1 % | 39,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 86 | [Qwen3-4B](https://quelleia.com/modeles/qwen3-4b.md) | Alibaba | Élevée | 37 % | 38,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 87 | [Grok-3 mini](https://quelleia.com/modeles/grok-3-mini.md) | xAI | Élevée | 36,8 % | 38,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 88 | [gpt-oss-20b](https://quelleia.com/modeles/gpt-oss-20b.md) | OpenAI | Faible | 36,8 % | 38,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 89 | [deepseek-r1-0528-qwen3-8b](https://quelleia.com/modeles/deepseek-r1-0528-qwen3-8b.md) | DeepSeek | non précisée | 36,6 % | 38,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 90 | [Reka Flash 3](https://quelleia.com/modeles/reka-flash-3.md) | Reka AI | non précisée | 35,9 % | 36,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 91 | [o4-mini](https://quelleia.com/modeles/o4-mini.md) | OpenAI | non précisée | 35,4 % | 35,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 91 | [Qwen3-235B-A22B](https://quelleia.com/modeles/qwen3-235b-a22b.md) | Alibaba | non précisée | 35,4 % | 35,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 93 | [Luciole 8B Instruct 1.1](https://quelleia.com/modeles/luciole-8b-instruct-1-1.md) | OpenLLM-France | non précisée | 35,4 % | 35,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 94 | [QwQ-32B](https://quelleia.com/modeles/qwq-32b.md) | Alibaba | non précisée | 35 % | 34,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 95 | [Apertus 70B Instruct](https://quelleia.com/modeles/apertus-70b-instruct.md) | Swiss AI (EPFL, ETH Zurich) | non précisée | 34,5 % | 33,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 96 | [o3](https://quelleia.com/modeles/o3.md) | OpenAI | non précisée | 33,5 % | 30,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 97 | [nvidia-nemotron-nano-9b-v2](https://quelleia.com/modeles/nvidia-nemotron-nano-9b-v2.md) | NVIDIA | non précisée | 33,4 % | 30,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 98 | [Qwen3-235B-A22B-Thinking (Jul 2025)](https://quelleia.com/modeles/qwen3-235b-a22b-thinking-jul-2025.md) | Alibaba | non précisée | 33,1 % | 29,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 99 | [Falcon 2 11B](https://quelleia.com/modeles/falcon-2-11b.md) | Technology Innovation Institute | non précisée | 33,1 % | 29,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 100 | [DeepSeek-R1-Distill-Llama-70B](https://quelleia.com/modeles/deepseek-r1-distill-llama-70b.md) | DeepSeek | non précisée | 32,6 % | 28,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 101 | [Grok-2 (Dec 2024)](https://quelleia.com/modeles/grok-2-dec-2024.md) | xAI | non précisée | 32,4 % | 28,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 102 | [GPT-5 mini](https://quelleia.com/modeles/gpt-5-mini.md) | OpenAI | Élevée | 32,2 % | 27,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 103 | [Apertus 8B Instruct](https://quelleia.com/modeles/apertus-8b-instruct.md) | Swiss AI (EPFL | non précisée | 32,1 % | 27,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 104 | [Mistral Small 3](https://quelleia.com/modeles/mistral-small-3.md) | Mistral AI | non précisée | 32 % | 27,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 105 | [GPT-5 nano](https://quelleia.com/modeles/gpt-5-nano.md) | OpenAI | Élevée | 31,6 % | 26,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 106 | [Llama 2-70B](https://quelleia.com/modeles/llama-2-70b.md) | Meta AI | non précisée | 31,4 % | 25,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 107 | [OLMo 2 Furious 13B](https://quelleia.com/modeles/olmo-2-furious-13b.md) | Allen Institute for AI | non précisée | 31,4 % | 25,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 108 | [Llama 3-8B](https://quelleia.com/modeles/llama-3-8b.md) | Meta AI | non précisée | 30,2 % | 22,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 109 | [Mistral 7B v0.1](https://quelleia.com/modeles/mistral-7b-v0-1.md) | Mistral AI | non précisée | 27,6 % | 16,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 110 | [DeepSeek-R1-Distill-Qwen-32B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-32b.md) | DeepSeek | non précisée | 26,3 % | 12,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 111 | [Mistral 7B v0.3](https://quelleia.com/modeles/mistral-7b-v0-3.md) | Mistral AI | non précisée | 26,1 % | 12,1 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 112 | [DeepSeek-R1-Distill-Qwen-14B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-14b.md) | DeepSeek | non précisée | 25,4 % | 10,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 113 | [EuroLLM 9B Instruct](https://quelleia.com/modeles/eurollm-9b-instruct.md) | EuroLLM | non précisée | 25,3 % | 9,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 114 | [Qwen3-Next-80B-A3B Thinking](https://quelleia.com/modeles/qwen3-next-80b-a3b-thinking.md) | Alibaba | Élevée | 24 % | 6,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 115 | [Mistral 7B v0.2](https://quelleia.com/modeles/mistral-7b-v0-2.md) | Mistral AI | non précisée | 23,4 % | 4,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 116 | [Gemma 7B](https://quelleia.com/modeles/gemma-7b.md) | Google DeepMind | non précisée | 20,9 % | -3,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 117 | [Llama 2-13B](https://quelleia.com/modeles/llama-2-13b.md) | Meta AI | non précisée | 20,1 % | -6,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 118 | [Llama 2-7B](https://quelleia.com/modeles/llama-2-7b.md) | Meta AI | non précisée | 16,3 % | -19,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 119 | [Llama 3.2 3B](https://quelleia.com/modeles/llama-3-2-3b.md) | Meta AI | non précisée | 15,4 % | -22,8 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 120 | [Qwen2.5-1.5B](https://quelleia.com/modeles/qwen2-5-1-5b.md) | Alibaba | non précisée | 14,8 % | -25,6 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 121 | [DeepSeek-R1-Distill-Llama-8B](https://quelleia.com/modeles/deepseek-r1-distill-llama-8b.md) | DeepSeek | non précisée | 14,7 % | -25,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 122 | [Salamandra 7B Instruct](https://quelleia.com/modeles/salamandra-7b-instruct.md) | Barcelona Supercomputing Center | non précisée | 13 % | -33,3 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 123 | [DeepSeek-R1-Distill-Qwen-7B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-7b.md) | DeepSeek | non précisée | 12,5 % | -35,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 124 | [Gemma 3 1B](https://quelleia.com/modeles/gemma-3-1b.md) | Google DeepMind | non précisée | 12,1 % | -37,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 125 | [Qwen3-0.6B](https://quelleia.com/modeles/qwen3-0-6b.md) | Alibaba | non précisée | 10,4 % | -46,9 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 126 | [Qwen3.5-0.8B](https://quelleia.com/modeles/qwen3-5-0-8b.md) | Alibaba | non précisée | 9,6 % | -51,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 127 | [Phi-4 Mini](https://quelleia.com/modeles/phi-4-mini.md) | Microsoft | Élevée | 8 % | -62,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 128 | [Ministral 8B](https://quelleia.com/modeles/ministral-8b.md) | Mistral AI | non précisée | 7,8 % | -63,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 129 | [DeepSeek-R1-Distill-Qwen-1.5B](https://quelleia.com/modeles/deepseek-r1-distill-qwen-1-5b.md) | DeepSeek | non précisée | 4,1 % | -99,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 130 | [Phi-4-Reasoning-plus](https://quelleia.com/modeles/phi-4-reasoning-plus.md) | Microsoft | non précisée | 3,8 % | -103,5 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 131 | [Llama 3.2 1B](https://quelleia.com/modeles/llama-3-2-1b.md) | Meta AI | non précisée | 3,4 % | -110,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 132 | [Phi-2](https://quelleia.com/modeles/phi-2.md) | Microsoft | non précisée | 3,2 % | -113,0 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 133 | [Gemma 3 270M](https://quelleia.com/modeles/gemma-3-270m.md) | Google DeepMind | non précisée | 2,7 % | -121,7 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 134 | [ALIA-40b](https://quelleia.com/modeles/alia-40b.md) | Barcelona Supercomputing Center | non précisée | 2,2 % | -134,2 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 135 | [Gemma 2B](https://quelleia.com/modeles/gemma-2b.md) | Google DeepMind | non précisée | 0,8 % | -175,4 | https://euroeval.com/leaderboards/Monolingual/french/ |
| 136 | [Phi-4 Reasoning](https://quelleia.com/modeles/phi-4-reasoning.md) | Microsoft | Élevée | 0,1 % | -175,4 | https://euroeval.com/leaderboards/Monolingual/french/ |

## En bref

**Que mesure EuroEval, ScaLA (grammaire) ?** Le modèle doit dire si une phrase française est correcte ou si elle a été abîmée. Sur Quelle IA, il est rangé dans la tâche Français.

**Comment EuroEval, ScaLA (grammaire) est-il noté ?** Un score d'accord avec la bonne réponse, de 0 à 100, corrigé du hasard. Un modèle qui obtient 50 % a un score IA d'environ 67.

**Qui est en tête sur EuroEval, ScaLA (grammaire) ?** GLM-5.3-Flash (Z.ai (Zhipu AI)) est en tête d'EuroEval, ScaLA (grammaire) avec 78 %, dans l'édition du 28 septembre 2026. Suivent Gemini 3 Pro (75,6 %) et Gemma 4 31B IT (74,2 %). 136 modèles y sont mesurés.

**Quelles sont les limites d'EuroEval, ScaLA (grammaire) ?** Juge la détection des fautes. La qualité de ce que le modèle écrit lui-même n'est pas notée. Au 28 septembre 2026, le benchmark est actif.

**Combien pèse EuroEval, ScaLA (grammaire) dans le score IA ?** EuroEval, ScaLA (grammaire) compte pour 0,62 % du score général, dans l'édition du 28 septembre 2026. Il porte 6 % de la tâche Français (10 % du score général), que se partagent 10 benchmarks.

**Qui maintient EuroEval, ScaLA (grammaire) ?** EuroEval. Sa page de référence : euroeval.com/leaderboards/Monolingual/french.

## Les autres benchmarks de la tâche Français

- [Arena, questions en français](https://quelleia.com/benchmarks/arena_francais.md) : 224 modèles · 2,5 % du score
- [compar:IA](https://quelleia.com/benchmarks/comparia.md) : 114 modèles · 2,5 % du score
- [EuroEval, Allociné (sentiment)](https://quelleia.com/benchmarks/euroeval_allocine.md) : 134 modèles · 0,62 % du score, saturé
- [EuroEval, ELTeC (noms propres)](https://quelleia.com/benchmarks/euroeval_eltec.md) : 134 modèles · 0,62 % du score
- [EuroEval, FQuAD (compréhension)](https://quelleia.com/benchmarks/euroeval_fquad.md) : 134 modèles · 0,62 % du score
- [EuroEval, HellaSwag (bon sens)](https://quelleia.com/benchmarks/euroeval_hellaswag_fr.md) : 134 modèles · 0,62 % du score
- [EuroEval, OrangeSum (résumé)](https://quelleia.com/benchmarks/euroeval_orange_sum.md) : 69 modèles · 0,62 % du score
- [EuroEval, INCLUDE (connaissances)](https://quelleia.com/benchmarks/euroeval_include_fr.md) : 64 modèles · 0,62 % du score
- [EuroEval, MultiLoKo (connaissances locales)](https://quelleia.com/benchmarks/euroeval_multiloko_fr.md) : 64 modèles · 0,62 % du score

Source : Quelle IA, édition du 28 septembre 2026. https://quelleia.com
