Classement · édition du
Classement des modèles de langage, Raisonnement
Quel modèle raisonne le mieux ?
Comment lire ce tableauComment lire ?
Le score Raisonnement résume les résultats d'un modèle sur les benchmarks de raisonnement où il a été mesuré. L'échelle est celle du score général : 100 correspond au meilleur modèle au lancement du site, 50 à GPT-4o de fin 2024. La barre indique la marge d'erreur : deux modèles dont les barres se chevauchent sont ex æquo.
Dans l'édition du 28 septembre 2026, GPT-6 Astra (OpenAI) est premier en raisonnement avec un score Raisonnement de 103,5, ex æquo avec Claude Opus 5.5.
| Rang | Empreinte | Modèle | Étiquettes | Score et marge (50 à 110) | Couverture | Semaine | Prix, € / M jetons |
|---|---|---|---|---|---|---|---|
| EX ÆQUO · RANG 1L'écart entre ces modèles est plus petit que la marge d'erreur. | |||||||
| 1 | GPT-6 AstraNOUVEAUOpenAI | 103,599,9 à 107,1 | 67 % · 6 mesures | Entrée | 8,78 € / 44 € | ||
| 2 | Claude Opus 5.5NOUVEAUAnthropic | 100,195,9 à 104,4 | 56 % · 5 mesures | Entrée | 4 € / 20 € | ||
| Sans rang | Claude Sonnet 5.5PROVISOIREAnthropic | 97,894,9 à 100,7 | 11 % · 1 mesure | Entrée | 2 € / 10 € | ||
| GROUPE 2Ex æquo · 9 modèles | |||||||
| 3 | GPT-5.6 SolOpenAI | 96,293,8 à 98,6 | 89 % · 8 mesures | 3,51 € / 18 € | |||
| 4 | GPT-5.5 ProOpenAI | 96,091,5 à 100,4 | 67 % · 6 mesures | 26 € / 158 € | |||
| 5 | Claude Fable 5.1NOUVEAUAnthropic | 95,492,4 à 98,5 | 67 % · 6 mesures | Entrée | 8,78 € / 44 € | ||
| 6 | Claude Opus 5Anthropic | 95,292,1 à 98,3 | 78 % · 7 mesures | 4,39 € / 22 € | |||
| 7 | Claude Fable 5Anthropic | 94,491,6 à 97,2 | 89 % · 8 mesures | 8,78 € / 44 € | |||
| 8 | GPT-5.5OpenAI | 94,191,6 à 96,6 | 89 % · 8 mesures | 4,34 € / 26 € | |||
| 9 | Gemini 3.8 FlashNOUVEAUGoogle DeepMind | 93,990,2 à 97,5 | 44 % · 4 mesures | Entrée | 1,50 € / 7,50 € | ||
| 10 | GPT-5.4 ProOpenAI | 93,189,7 à 96,5 | 56 % · 5 mesures | 26 € / 158 € | |||
| Sans rang | Gemini 3 Deep ThinkPROVISOIREGoogle DeepMind | 92,689,0 à 96,1 | 22 % · 2 mesures | Inconnu | |||
| 11 | Gemini 3.1 ProGoogle DeepMind | 91,888,5 à 95,0 | 100 % · 9 mesures | 2 € / 12 € | |||
| GROUPE 3Ex æquo · 14 modèles | |||||||
| 12 | Gemini 3.7 FlashGoogle DeepMind | 91,388,9 à 93,6 | 67 % · 6 mesures | 1,50 € / 7,50 € | |||
| Sans rang | Muse Spark 1.2PROVISOIREMeta AI | 90,785,2 à 96,3 | 33 % · 3 mesures | 1,25 € / 4,25 € | |||
| 13 | GPT-5.6 TerraOpenAI | 90,586,8 à 94,1 | 78 % · 7 mesures | 2,17 € / 13 € | |||
| 14 | Gemini 3.5 FlashGoogle DeepMind | 89,787,5 à 91,9 | 100 % · 9 mesures | 1,30 € / 7,80 € | |||
| 15 | GPT-5.4OpenAI | 89,487,6 à 91,3 | 89 % · 8 mesures | 2,17 € / 13 € | |||
| 16 | Claude Opus 4.8Anthropic | 89,287,3 à 91,0 | 100 % · 9 mesures | 4,39 € / 22 € | |||
| Sans rang | Qwen3.8 Max (0902)PROVISOIREAlibaba | 88,984,4 à 93,5 | 11 % · 1 mesure | Entrée | 1,76 € / 5,27 € | ||
| 17 | DeepSeek V4 Pro 0813DeepSeek | 88,685,8 à 91,5 | 67 % · 6 mesures | 0,61 € / 1,82 € | |||
| 18 | Grok 4.6xAI | 88,686,4 à 90,9 | 78 % · 7 mesures | 1,76 € / 5,27 € | |||
| Sans rang | Muse Spark 1.1PROVISOIREMeta AI | 88,682,4 à 94,9 | 22 % · 2 mesures | 1,10 € / 3,73 € | |||
| 19 | Claude Opus 4.7Anthropic | 88,485,8 à 91,0 | 89 % · 8 mesures | 4,39 € / 22 € | |||
| 20 | Claude Sonnet 5Anthropic | 87,684,6 à 90,6 | 67 % · 6 mesures | 1,73 € / 8,67 € | |||
| 21 | Muse Spark 1.3NOUVEAUMeta AI | 87,683,8 à 91,4 | 44 % · 4 mesures | Entrée | 1,25 € / 4,25 € | ||
| 22 | Qwen 3.8 MaxAlibaba | 87,683,6 à 91,6 | 44 % · 4 mesures | 1,76 € / 5,27 € | |||
| 23 | Gemini 3.6 FlashGoogle DeepMind | 87,385,3 à 89,4 | 67 % · 6 mesures | 1,30 € / 6,50 € | |||
| 24 | Kimi K3Moonshot | 87,185,2 à 89,0 | 89 % · 8 mesures | 2,60 € / 13 € | |||
| 25 | Claude Opus 4.6Anthropic | 86,883,7 à 89,9 | 100 % · 9 mesures | 4,39 € / 22 € | |||
| GROUPE 4Ex æquo · 10 modèles | |||||||
| 26 | DeepSeek V4 Flash 0731DeepSeek | 86,884,8 à 88,8 | 78 % · 7 mesures | 0,18 € / 0,35 € | |||
| 27 | Grok 4.20xAI | 86,382,5 à 90,0 | 67 % · 6 mesures | 1,08 € / 2,17 € | |||
| 28 | GLM-5.3Z.ai (Zhipu AI) | 86,080,8 à 91,2 | 44 % · 4 mesures | 1,40 € / 4,40 € | |||
| 29 | Grok 4.5xAI | 86,083,8 à 88,3 | 67 % · 6 mesures | 1,76 € / 5,27 € | |||
| 30 | GPT-5.6 LunaOpenAI | 85,883,0 à 88,5 | 78 % · 7 mesures | 0,87 € / 5,20 € | |||
| 31 | Qwen3.7-MaxAlibaba | 85,580,9 à 90,1 | 56 % · 5 mesures | 1,08 € / 3,25 € | |||
| Sans rang | DeepSeek V4.1 FlashPROVISOIREDeepSeek | 85,580,0 à 91,1 | 22 % · 2 mesures | Entrée | 0,15 € / 0,60 € | ||
| Sans rang | GPT-5.2 ProPROVISOIREOpenAI | 85,583,1 à 87,9 | 33 % · 3 mesures | 18 € / 147 € | |||
| 32 | Claude Sonnet 4.6Anthropic | 85,081,3 à 88,7 | 78 % · 7 mesures | 2,63 € / 13 € | |||
| 33 | GPT-5.2OpenAI | 85,081,9 à 88,0 | 100 % · 9 mesures | 1,52 € / 12 € | |||
| 34 | Gemini 3 FlashGoogle DeepMind | 83,280,8 à 85,6 | 89 % · 8 mesures | 0,43 € / 2,60 € | |||
| 35 | Gemini 3 ProGoogle DeepMind | 82,678,6 à 86,7 | 67 % · 6 mesures | 1,73 € / 10 € | |||
| GROUPE 5Ex æquo · 17 modèles | |||||||
| 36 | Claude Opus 4.5Anthropic | 82,480,3 à 84,5 | 100 % · 9 mesures | 4,39 € / 22 € | |||
| 37 | Inkling-SmallThinking Machines | 82,478,4 à 86,3 | 44 % · 4 mesures | 0,45 € / 1,20 € | |||
| 38 | Qwen 3.6 Max (Preview)Alibaba | 81,677,8 à 85,4 | 56 % · 5 mesures | 1,14 € / 6,85 € | |||
| Sans rang | Qwen 3.8 27BPROVISOIREAlibaba | 81,676,3 à 86,9 | 22 % · 2 mesures | 0,35 € / 2,75 € | |||
| 39 | InklingThinking Machines | 81,378,8 à 83,8 | 67 % · 6 mesures | 0,95 € / 4,05 € | |||
| 40 | Kimi K2.6Moonshot | 81,377,4 à 85,3 | 44 % · 4 mesures | 0,59 € / 2,96 € | |||
| 41 | Grok 4.3 BetaxAI | 81,176,4 à 85,8 | 44 % · 4 mesures | 1,08 € / 2,17 € | |||
| 42 | GLM-5.2Z.ai (Zhipu AI) | 80,878,5 à 83,1 | 78 % · 7 mesures | 0,98 € / 3,08 € | |||
| 43 | DeepSeek-V4-ProDeepSeek | 80,676,3 à 84,8 | 56 % · 5 mesures | 0,38 € / 0,75 € | |||
| Sans rang | Kimi K2.7 CodePROVISOIREMoonshot | 80,675,9 à 85,2 | 22 % · 2 mesures | 0,83 € / 3,51 € | |||
| Sans rang | Grok 4.1 FastPROVISOIRExAI | 80,375,3 à 85,3 | 33 % · 3 mesures | 0,17 € / 0,43 € | |||
| 44 | GPT-5.1OpenAI | 79,877,3 à 82,2 | 100 % · 9 mesures | 1,08 € / 8,67 € | |||
| 45 | GPT-5 ProOpenAI | 79,576,3 à 82,7 | 44 % · 4 mesures | 13 € / 105 € | |||
| 46 | GPT-5OpenAI | 79,075,5 à 82,4 | 100 % · 9 mesures | 1,08 € / 8,67 € | |||
| Sans rang | GLM-5.1PROVISOIREZ.ai (Zhipu AI) | 79,074,2 à 83,8 | 22 % · 2 mesures | 0,85 € / 2,67 € | |||
| 47 | Grok 4xAI | 78,775,6 à 81,9 | 56 % · 5 mesures | 2,63 € / 13 € | |||
| 48 | Qwen3.7-PlusAlibaba | 78,774,5 à 82,9 | 44 % · 4 mesures | 0,28 € / 1,11 € | |||
| Sans rang | Qwen3.7 FlashPROVISOIREAlibaba | 78,572,6 à 84,3 | 22 % · 2 mesures | 0,03 € / 0,11 € | |||
| 49 | Nemotron 3 UltraNVIDIA | 78,272,4 à 84,0 | 44 % · 4 mesures | Inconnu | |||
| 50 | o3OpenAI | 77,973,8 à 82,1 | 100 % · 9 mesures | 1,76 € / 7,02 € | |||
| Sans rang | DeepSeek-V3.2-SpecialePROVISOIREDeepSeek | 77,972,3 à 83,5 | 11 % · 1 mesure | 0,51 € / 1,47 € | |||
| 51 | GPT-5.4 MiniOpenAI | 77,775,1 à 80,3 | 78 % · 7 mesures | 0,65 € / 3,90 € | |||
| 52 | Qwen3.5 397B-A17BAlibaba | 77,772,4 à 83,0 | 44 % · 4 mesures | 0,34 € / 2,03 € | |||
| GROUPE 6Ex æquo · 21 modèles | |||||||
| 53 | Claude Sonnet 4.5Anthropic | 77,275,0 à 79,3 | 100 % · 9 mesures | 2,60 € / 13 € | |||
| 54 | Claude Opus 4.1Anthropic | 76,971,2 à 82,6 | 78 % · 7 mesures | 13 € / 66 € | |||
| Sans rang | DeepSeek-V4-FlashPROVISOIREDeepSeek | 76,672,1 à 81,2 | 33 % · 3 mesures | 0,09 € / 0,17 € | |||
| Sans rang | Kimi K2 ThinkingPROVISOIREMoonshot | 76,669,6 à 83,6 | 11 % · 1 mesure | 0,52 € / 2,17 € | |||
| 55 | Kimi K2.5Moonshot | 76,473,2 à 79,5 | 56 % · 5 mesures | 0,35 € / 1,65 € | |||
| 56 | Qwen 3.5 Plus (hosted 397B-A17B)Alibaba | 76,471,4 à 81,3 | 44 % · 4 mesures | 0,35 € / 2,11 € | |||
| 57 | o3-proOpenAI | 76,173,4 à 78,8 | 44 % · 4 mesures | 18 € / 70 € | |||
| Sans rang | MiniMax-M2.5PROVISOIREMiniMax | 76,172,2 à 80,0 | 22 % · 2 mesures | 0,13 € / 1 € | |||
| Sans rang | Qwen3.5-122B-A10BPROVISOIREAlibaba | 75,970,7 à 81,0 | 33 % · 3 mesures | 0,35 € / 2,81 € | |||
| 58 | Gemini 3.5 Flash-LiteGoogle DeepMind | 75,672,9 à 78,2 | 67 % · 6 mesures | 0,26 € / 2,17 € | |||
| Sans rang | DeepSeek-V3.2-ExpPROVISOIREDeepSeek | 75,668,4 à 82,8 | 22 % · 2 mesures | 0,25 € / 0,38 € | |||
| 59 | GPT-5 miniOpenAI | 75,172,4 à 77,8 | 89 % · 8 mesures | 0,22 € / 1,73 € | |||
| 60 | Qwen 3.5 Flash (hosted 35B-A3B)Alibaba | 75,169,7 à 80,5 | 44 % · 4 mesures | 0,09 € / 0,35 € | |||
| Sans rang | MiMo-V2.5-ProPROVISOIREXiaomi | 75,169,4 à 80,8 | 22 % · 2 mesures | 0,38 € / 0,76 € | |||
| 61 | DeepSeek-V3.2DeepSeek | 74,872,1 à 77,5 | 56 % · 5 mesures | 0,20 € / 0,30 € | |||
| 62 | GPT-5.4 NanoOpenAI | 74,871,8 à 77,8 | 78 % · 7 mesures | 0,17 € / 1,08 € | |||
| 63 | Qwen 3.6 PlusAlibaba | 74,870,4 à 79,2 | 44 % · 4 mesures | 0,28 € / 1,69 € | |||
| Sans rang | Qwen3.5-27BPROVISOIREAlibaba | 74,869,8 à 79,8 | 22 % · 2 mesures | 0,26 € / 2,11 € | |||
| Sans rang | GLM-5.3-FlashPROVISOIREZ.ai (Zhipu AI) | 74,567,5 à 81,6 | 22 % · 2 mesures | 0,15 € / 0,50 € | |||
| 64 | o4-miniOpenAI | 74,370,5 à 78,1 | 100 % · 9 mesures | 0,95 € / 3,82 € | |||
| Sans rang | Gemma 4 31B ITPROVISOIREGoogle DeepMind | 74,367,0 à 81,6 | 33 % · 3 mesures | 0,10 € / 0,31 € | |||
| 65 | Gemini 2.5 Pro (Jun 2025)Google DeepMind | 74,070,1 à 77,9 | 89 % · 8 mesures | 1,10 € / 8,78 € | |||
| 66 | Qwen3.6 27BAlibaba | 74,068,5 à 79,5 | 44 % · 4 mesures | 0,53 € / 3,16 € | |||
| 67 | Grok 4 FastxAI | 73,571,1 à 75,9 | 44 % · 4 mesures | 0,17 € / 0,43 € | |||
| 68 | GLM-5Z.ai (Zhipu AI) | 73,269,7 à 76,8 | 56 % · 5 mesures | 0,52 € / 1,66 € | |||
| 69 | Gemini 3.1 Flash-LiteGoogle DeepMind | 73,266,5 à 80,0 | 56 % · 5 mesures | 0,22 € / 1,30 € | |||
| 70 | MiniMax-M3MiniMax | 73,269,3 à 77,2 | 67 % · 6 mesures | 0,24 € / 0,97 € | |||
| 71 | Claude Opus 4Anthropic | 73,068,7 à 77,3 | 78 % · 7 mesures | 13 € / 66 € | |||
| 72 | Claude Haiku 4.5Anthropic | 72,268,0 à 76,4 | 67 % · 6 mesures | 0,88 € / 4,39 € | |||
| Sans rang | DeepSeek-V3.1-TerminusPROVISOIREDeepSeek | 72,267,2 à 77,2 | 22 % · 2 mesures | 0,24 € / 0,88 € | |||
| Sans rang | GLM-4.7PROVISOIREZ.ai (Zhipu AI) | 72,264,5 à 79,9 | 22 % · 2 mesures | 0,35 € / 1,52 € | |||
| 73 | Qwen 3.6 35B-A3BAlibaba | 71,963,3 à 80,5 | 44 % · 4 mesures | 0,22 € / 1,30 € | |||
| Sans rang | GPT-5.5 InstantPROVISOIREOpenAI | 71,963,9 à 79,9 | 11 % · 1 mesure | 4,39 € / 26 € | |||
| Sans rang | tiny-recursion-modelPROVISOIREÉditeur inconnu | 71,768,9 à 74,5 | 22 % · 2 mesures | Inconnu | |||
| GROUPE 7Ex æquo · 13 modèles | |||||||
| 74 | Claude Sonnet 4Anthropic | 71,469,0 à 73,8 | 78 % · 7 mesures | 2,60 € / 13 € | |||
| 75 | Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba | 71,466,7 à 76,1 | 44 % · 4 mesures | 0,26 € / 2,55 € | |||
| Sans rang | Kimi K2 (Sep 2025)PROVISOIREMoonshot | 71,267,1 à 75,2 | 11 % · 1 mesure | 0,55 € / 2,22 € | |||
| 76 | Gemini 2.5 Pro (Mar 2025)Google DeepMind | 70,965,9 à 75,9 | 44 % · 4 mesures | 2,19 € / 8,78 € | |||
| 77 | Qwen3-MaxAlibaba | 70,965,0 à 76,8 | 44 % · 4 mesures | 0,68 € / 3,38 € | |||
| Sans rang | Qwen3.5-35B-A3BPROVISOIREAlibaba | 70,965,9 à 75,9 | 33 % · 3 mesures | 0,12 € / 0,87 € | |||
| 78 | DeepSeek-V3.1DeepSeek | 70,665,4 à 75,9 | 44 % · 4 mesures | 0,18 € / 0,69 € | |||
| Sans rang | Qwen Plus (Apr 2025)PROVISOIREAlibaba | 70,464,3 à 76,4 | 22 % · 2 mesures | Inconnu | |||
| 79 | Qwen 3.6 FlashAlibaba | 70,164,6 à 75,6 | 56 % · 5 mesures | 0,16 € / 0,99 € | |||
| Sans rang | Gemini 2.5 Flash (May 2025)PROVISOIREGoogle DeepMind | 69,866,8 à 72,9 | 33 % · 3 mesures | 0,12 € / 2,77 € | |||
| Sans rang | Gemini 2.5 Flash (Apr 2025)PROVISOIREGoogle DeepMind | 69,666,5 à 72,7 | 22 % · 2 mesures | 0,13 € / 0,53 € | |||
| 80 | Claude 3.7 SonnetAnthropic | 69,365,5 à 73,1 | 56 % · 5 mesures | 2,63 € / 13 € | |||
| 81 | Gemini 2.5 Flash (Jun 2025)Google DeepMind | 69,365,2 à 73,5 | 44 % · 4 mesures | 0,26 € / 2,17 € | |||
| Sans rang | seed-oss-36b-instructPROVISOIREByteDance | 68,860,6 à 77,0 | 11 % · 1 mesure | Inconnu | |||
| 82 | o1OpenAI | 68,565,7 à 71,4 | 67 % · 6 mesures | 13 € / 53 € | |||
| Sans rang | codex-mini-2025-05-16PROVISOIREOpenAI | 68,365,0 à 71,5 | 22 % · 2 mesures | 1,32 € / 5,27 € | |||
| Sans rang | Gemma 4 26B A4BPROVISOIREGoogle DeepMind | 68,362,7 à 73,8 | 33 % · 3 mesures | 0,05 € / 0,29 € | |||
| Sans rang | Mistral Medium 3.5PROVISOIREMistral AI | 68,063,1 à 72,9 | 22 % · 2 mesures | 1,30 € / 6,50 € | |||
| Sans rang | DeepSeek-R1 (May 2025)PROVISOIREDeepSeek | 67,063,2 à 70,7 | 33 % · 3 mesures | 0,43 € / 1,86 € | |||
| Sans rang | o1-proPROVISOIREOpenAI | 67,062,9 à 71,1 | 22 % · 2 mesures | 132 € / 527 € | |||
| 83 | o3-miniOpenAI | 66,259,2 à 73,2 | 89 % · 8 mesures | 0,97 € / 3,86 € | |||
| 84 | Qwen3-235B-A22BAlibaba | 65,960,9 à 71,0 | 44 % · 4 mesures | 0,61 € / 2,46 € | |||
| Sans rang | o1-previewPROVISOIREOpenAI | 65,961,5 à 70,4 | 22 % · 2 mesures | Inconnu | |||
| 85 | Qwen3-235B-A22B-Instruct (Jul 2025)Alibaba | 65,160,0 à 70,3 | 44 % · 4 mesures | 0,19 € / 0,77 € | |||
| Sans rang | Qwen3.5-9BPROVISOIREAlibaba | 64,959,6 à 70,2 | 33 % · 3 mesures | 0,10 € / 0,15 € | |||
| Sans rang | Grok-3 miniPROVISOIRExAI | 64,660,2 à 69,0 | 22 % · 2 mesures | 0,26 € / 0,43 € | |||
| Sans rang | Gemini 2.5 Pro (May 2025)PROVISOIREGoogle DeepMind | 64,459,0 à 69,7 | 11 % · 1 mesure | 2,19 € / 8,78 € | |||
| 86 | gpt-oss-120bOpenAI | 64,155,6 à 72,6 | 56 % · 5 mesures | 0,03 € / 0,16 € | |||
| GROUPE 8Ex æquo · 10 modèles | |||||||
| 87 | DeepSeek-R1DeepSeek | 63,358,5 à 68,2 | 44 % · 4 mesures | Inconnu | |||
| Sans rang | Mistral Small 4PROVISOIREMistral AI | 63,158,2 à 68,0 | 22 % · 2 mesures | 0,13 € / 0,52 € | |||
| 88 | Claude 3.5 Sonnet (October 2024)Anthropic | 62,855,5 à 70,1 | 44 % · 4 mesures | Inconnu | |||
| 89 | GPT-4.5OpenAI | 62,557,1 à 67,9 | 56 % · 5 mesures | Inconnu | |||
| Sans rang | GLM-4.5-AirPROVISOIREZ.ai (Zhipu AI) | 62,558,7 à 66,4 | 11 % · 1 mesure | 0,18 € / 0,97 € | |||
| Sans rang | Gemini 2.0 ProPROVISOIREGoogle DeepMind | 62,357,7 à 66,8 | 11 % · 1 mesure | 1,75 € / 6,98 € | |||
| Sans rang | Qwen3-30B-A3B-Thinking (Jul 2025)PROVISOIREAlibaba | 62,057,1 à 66,9 | 33 % · 3 mesures | 0,32 € / 1,14 € | |||
| Sans rang | Grok 3PROVISOIRExAI | 61,254,7 à 67,8 | 33 % · 3 mesures | 2,63 € / 13 € | |||
| Sans rang | Qwen3-30B-A3B-Instruct (Jul 2025)PROVISOIREAlibaba | 61,055,4 à 66,6 | 33 % · 3 mesures | 0,26 € / 0,44 € | |||
| Sans rang | Mistral Medium 3.1PROVISOIREMistral AI | 60,755,6 à 65,8 | 22 % · 2 mesures | 0,35 € / 1,73 € | |||
| 90 | GPT-4.1OpenAI | 60,555,7 à 65,2 | 89 % · 8 mesures | 1,76 € / 7,02 € | |||
| Sans rang | Gemini 2.0 Flash Thinking (Jan 2025)PROVISOIREGoogle DeepMind | 60,552,9 à 68,0 | 22 % · 2 mesures | Inconnu | |||
| 91 | GPT-4.1 miniOpenAI | 60,255,3 à 65,0 | 67 % · 6 mesures | 0,35 € / 1,39 € | |||
| Sans rang | DeepSeek-R1-Distill-Qwen-32BPROVISOIREDeepSeek | 60,253,9 à 66,4 | 11 % · 1 mesure | 0,25 € / 0,76 € | |||
| Sans rang | Gemini 2.0 Flash (Dec 2024)PROVISOIREGoogle DeepMind | 60,252,6 à 67,8 | 11 % · 1 mesure | 1,10 € / 4,39 € | |||
| Sans rang | glm-4.7-flashPROVISOIREZ.ai (Zhipu AI) | 59,953,8 à 66,1 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Qwen3-32BPROVISOIREAlibaba | 59,954,9 à 64,9 | 33 % · 3 mesures | 0,07 € / 0,24 € | |||
| 92 | GPT-5 nanoOpenAI | 59,750,7 à 68,6 | 78 % · 7 mesures | 0,04 € / 0,35 € | |||
| Sans rang | QwQ-32BPROVISOIREAlibaba | 59,753,6 à 65,7 | 11 % · 1 mesure | 0,13 € / 0,35 € | |||
| Sans rang | Kimi K2 (Jul 2025)PROVISOIREMoonshot | 59,447,8 à 71,0 | 22 % · 2 mesures | 0,50 € / 2,02 € | |||
| Sans rang | gpt-oss-20bPROVISOIREOpenAI | 59,153,7 à 64,6 | 33 % · 3 mesures | 0,03 € / 0,12 € | |||
| Sans rang | Claude 3.5 SonnetPROVISOIREAnthropic | 58,953,3 à 64,4 | 33 % · 3 mesures | 2,63 € / 13 € | |||
| Sans rang | Mistral Large 3PROVISOIREMistral AI | 58,653,4 à 63,9 | 22 % · 2 mesures | 0,43 € / 1,30 € | |||
| Sans rang | Qwen3-14BPROVISOIREAlibaba | 57,852,2 à 63,4 | 33 % · 3 mesures | 0,31 € / 1,23 € | |||
| Sans rang | DeepSeek-V3 (Mar 2025)PROVISOIREDeepSeek | 57,352,5 à 62,1 | 33 % · 3 mesures | 0,17 € / 0,67 € | |||
| Sans rang | Gemini 2.5 Flash-Lite (Jun 2025)PROVISOIREGoogle DeepMind | 57,151,1 à 63,1 | 22 % · 2 mesures | 0,09 € / 0,35 € | |||
| Sans rang | GPT-4o (May 2024)PROVISOIREOpenAI | 56,851,5 à 62,1 | 33 % · 3 mesures | 4,39 € / 13 € | |||
| Sans rang | Magistral Small 1.0PROVISOIREMistral AI | 56,848,7 à 64,9 | 22 % · 2 mesures | 0,43 € / 1,30 € | |||
| Sans rang | qwen3-4b-instruct-2507PROVISOIREAlibaba | 56,851,8 à 61,8 | 11 % · 1 mesure | 0,18 € / 0,18 € | |||
| Sans rang | Gemini 2.0 Flash (Feb 2025)PROVISOIREGoogle DeepMind | 56,049,9 à 62,1 | 33 % · 3 mesures | 0,09 € / 0,37 € | |||
| Sans rang | DeepSeek-R1-Distill-Qwen-14BPROVISOIREDeepSeek | 55,851,1 à 60,4 | 11 % · 1 mesure | 0,13 € / 0,38 € | |||
| Sans rang | Grok-2 (Dec 2024)PROVISOIRExAI | 55,548,2 à 62,8 | 22 % · 2 mesures | Inconnu | |||
| 93 | Llama 4 MaverickMeta AI | 55,250,3 à 60,2 | 78 % · 7 mesures | 0,13 € / 0,52 € | |||
| Sans rang | Mistral Medium 3PROVISOIREMistral AI | 55,249,4 à 61,0 | 11 % · 1 mesure | 0,35 € / 1,76 € | |||
| Sans rang | o1-miniPROVISOIREOpenAI | 55,042,3 à 67,7 | 33 % · 3 mesures | 0,97 € / 3,86 € | |||
| Sans rang | Qwen2.5-72BPROVISOIREAlibaba | 54,448,7 à 60,2 | 33 % · 3 mesures | 2,46 € / 7,37 € | |||
| Sans rang | Qwen3-30B-A3BPROVISOIREAlibaba | 54,448,3 à 60,6 | 33 % · 3 mesures | 0,10 € / 0,43 € | |||
| 94 | GPT-4 (Jun 2023)OpenAI | 54,246,9 à 61,5 | 56 % · 5 mesures | Inconnu | |||
| 95 | Claude 3 OpusAnthropic | 53,747,9 à 59,4 | 67 % · 6 mesures | 13 € / 66 € | |||
| Sans rang | Magistral Small 1.2PROVISOIREMistral AI | 53,747,8 à 59,6 | 22 % · 2 mesures | 0,44 € / 1,32 € | |||
| Sans rang | Cohere Command APROVISOIRECohere | 53,447,7 à 59,1 | 22 % · 2 mesures | 2,17 € / 8,67 € | |||
| Sans rang | Magistral Medium 1.1PROVISOIREMistral AI | 53,445,0 à 61,8 | 22 % · 2 mesures | 1,73 € / 4,34 € | |||
| Sans rang | Qwen3-4BPROVISOIREAlibaba | 53,449,6 à 57,2 | 11 % · 1 mesure | 0,03 € / 0,03 € | |||
| Sans rang | Llama 3.1-405BPROVISOIREMeta AI | 52,946,8 à 58,9 | 33 % · 3 mesures | 2,38 € / 2,60 € | |||
| Sans rang | GPT-4 Turbo (Nov 2023)PROVISOIREOpenAI | 52,646,2 à 59,1 | 22 % · 2 mesures | 8,78 € / 26 € | |||
| Sans rang | Mistral Large 2 (Nov 2024)PROVISOIREMistral AI | 52,646,1 à 59,1 | 22 % · 2 mesures | 1,76 € / 5,27 € | |||
| Sans rang | Mistral Small 3.2PROVISOIREMistral AI | 52,646,4 à 58,8 | 22 % · 2 mesures | 0,07 € / 0,17 € | |||
| Sans rang | Gemini 1.5 Pro (Sept 2024)PROVISOIREGoogle DeepMind | 52,446,2 à 58,5 | 33 % · 3 mesures | Inconnu | |||
| Sans rang | Llama 3.1-70BPROVISOIREMeta AI | 52,446,5 à 58,2 | 22 % · 2 mesures | 0,63 € / 0,63 € | |||
| Sans rang | Mistral Large 2 (Jul 2024)PROVISOIREMistral AI | 52,145,4 à 58,8 | 33 % · 3 mesures | 1,76 € / 5,27 € | |||
| Sans rang | Qwen3-8BPROVISOIREAlibaba | 52,146,2 à 58,0 | 33 % · 3 mesures | 0,04 € / 0,35 € | |||
| 96 | Llama 3.3 70BMeta AI | 51,845,1 à 58,6 | 44 % · 4 mesures | 0,09 € / 0,28 € | |||
| Sans rang | GPT-4o (Nov 2024)PROVISOIREOpenAI | 51,844,5 à 59,1 | 33 % · 3 mesures | 2,19 € / 8,78 € | |||
| Sans rang | glm-4.7-flash_nonePROVISOIREZ.ai (Zhipu AI) | 51,648,3 à 54,8 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | GPT-4 Turbo (Apr 2024)PROVISOIREOpenAI | 51,343,0 à 59,6 | 33 % · 3 mesures | Inconnu | |||
| Sans rang | Mistral Small 3.1PROVISOIREMistral AI | 50,544,3 à 56,8 | 22 % · 2 mesures | 0,09 € / 0,26 € | |||
| GROUPE 9Ex æquo · 3 modèles | |||||||
| 97 | Llama 4 ScoutMeta AI | 50,344,4 à 56,1 | 56 % · 5 mesures | 0,09 € / 0,26 € | |||
| Sans rang | Claude 3.5 HaikuPROVISOIREAnthropic | 48,742,3 à 55,1 | 11 % · 1 mesure | 0,70 € / 3,51 € | |||
| Sans rang | DeepSeek-V3PROVISOIREDeepSeek | 48,738,8 à 58,6 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Gemini 1.5 Pro (May 2024)PROVISOIREGoogle DeepMind | 48,742,1 à 55,3 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Qwen2.5-32BPROVISOIREAlibaba | 48,446,0 à 50,9 | 11 % · 1 mesure | 0,61 € / 2,46 € | |||
| Sans rang | Gemma 3 27BPROVISOIREGoogle DeepMind | 46,638,3 à 55,0 | 33 % · 3 mesures | 0,07 € / 0,14 € | |||
| Sans rang | GPT-4o (Aug 2024)PROVISOIREOpenAI | 46,435,1 à 57,6 | 22 % · 2 mesures | 2,19 € / 8,78 € | |||
| Sans rang | Gemini 1.5 Flash (Sep 2024)PROVISOIREGoogle DeepMind | 45,839,0 à 52,7 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Llama 3.2 90BPROVISOIREMeta AI | 45,844,8 à 46,8 | 11 % · 1 mesure | 0,31 € / 0,35 € | |||
| Sans rang | Mistral LargePROVISOIREMistral AI | 45,839,0 à 52,7 | 11 % · 1 mesure | 0,44 € / 1,32 € | |||
| Sans rang | Pixtral LargePROVISOIREMistral AI | 45,844,8 à 46,8 | 11 % · 1 mesure | 1,76 € / 5,27 € | |||
| Sans rang | Gemini 2.0 Flash-LitePROVISOIREGoogle DeepMind | 45,636,8 à 54,3 | 22 % · 2 mesures | 0,07 € / 0,26 € | |||
| 98 | GPT-4o miniOpenAI | 45,038,0 à 52,1 | 67 % · 6 mesures | 0,13 € / 0,53 € | |||
| Sans rang | Llama 3-70BPROVISOIREMeta AI | 45,038,0 à 52,1 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Phi-4PROVISOIREMicrosoft | 45,043,3 à 46,8 | 11 % · 1 mesure | 0,06 € / 0,12 € | |||
| 99 | GPT-4.1 nanoOpenAI | 44,537,3 à 51,7 | 44 % · 4 mesures | 0,09 € / 0,35 € | |||
| Sans rang | Mistral Small 3PROVISOIREMistral AI | 44,542,8 à 46,2 | 11 % · 1 mesure | 0,04 € / 0,07 € | |||
| Sans rang | Claude 3 SonnetPROVISOIREAnthropic | 44,337,2 à 51,3 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Command R+PROVISOIRECohere | 44,036,2 à 51,8 | 33 % · 3 mesures | 2,19 € / 8,78 € | |||
| Sans rang | mistral-small-2402PROVISOIREMistral AI | 44,036,9 à 51,1 | 11 % · 1 mesure | 0,13 € / 0,53 € | |||
| Sans rang | Mixtral 8x22BPROVISOIREMistral AI | 43,533,4 à 53,6 | 22 % · 2 mesures | 1,76 € / 5,27 € | |||
| Sans rang | Claude 2PROVISOIREAnthropic | 41,734,2 à 49,1 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Gemini 1.5 Flash (May 2024)PROVISOIREGoogle DeepMind | 41,433,8 à 49,0 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Mistral MediumPROVISOIREMistral AI | 41,133,6 à 48,6 | 11 % · 1 mesure | 1,32 € / 6,58 € | |||
| Sans rang | Gemma 3 12BPROVISOIREGoogle DeepMind | 40,332,1 à 48,6 | 33 % · 3 mesures | 0,04 € / 0,13 € | |||
| Sans rang | Gemma 3 4BPROVISOIREGoogle DeepMind | 39,631,6 à 47,5 | 33 % · 3 mesures | 0,04 € / 0,09 € | |||
| Sans rang | Claude 2.1PROVISOIREAnthropic | 39,330,8 à 47,8 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Ministral 3BPROVISOIREMistral AI | 39,331,0 à 47,6 | 22 % · 2 mesures | 0,04 € / 0,04 € | |||
| Sans rang | Claude 3 HaikuPROVISOIREAnthropic | 39,030,8 à 47,3 | 33 % · 3 mesures | 0,22 € / 1,10 € | |||
| Sans rang | Gemini 1.5 Flash 8BPROVISOIREGoogle DeepMind | 38,831,0 à 46,6 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Mistral Small v24.09PROVISOIREMistral AI | 38,831,0 à 46,6 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Gemma 2 27BPROVISOIREGoogle DeepMind | 38,530,2 à 46,9 | 22 % · 2 mesures | 0,57 € / 0,57 € | |||
| Sans rang | Llama 3.1-8BPROVISOIREMeta AI | 36,426,9 à 46,0 | 33 % · 3 mesures | 0,02 € / 0,03 € | |||
| Sans rang | Mistral NeMoPROVISOIREMistral AI | 36,428,4 à 44,5 | 11 % · 1 mesure | 0,13 € / 0,13 € | |||
| Sans rang | Qwen2.5-7BPROVISOIREAlibaba | 36,228,1 à 44,3 | 33 % · 3 mesures | 0,15 € / 0,61 € | |||
| Sans rang | c4ai-command-r-08-2024PROVISOIRECohere | 35,926,2 à 45,6 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Mixtral 8x7BPROVISOIREMistral AI | 34,124,6 à 43,6 | 22 % · 2 mesures | 0,61 € / 0,61 € | |||
| GROUPE 10Un seul modèle dans ce groupe | |||||||
| 100 | GPT-3.5 Turbo (Jan 2024)OpenAI | 33,823,9 à 43,8 | 56 % · 5 mesures | 0,44 € / 1,32 € | |||
| Sans rang | deepseek-r1-0528-qwen3-8bPROVISOIREDeepSeek | 33,633,0 à 34,2 | 11 % · 1 mesure | 0,05 € / 0,08 € | |||
| Sans rang | granite-4.0-microPROVISOIREIBM | 33,332,7 à 33,9 | 11 % · 1 mesure | 0,01 € / 0,10 € | |||
| Sans rang | Qwen3-1.7BPROVISOIREAlibaba | 33,032,5 à 33,6 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Gemini 1.0 ProPROVISOIREGoogle DeepMind | 32,824,6 à 41,0 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Qwen1.5-110BPROVISOIREAlibaba | 32,329,2 à 35,3 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Ministral 8BPROVISOIREMistral AI | 31,523,3 à 39,7 | 11 % · 1 mesure | 0,09 € / 0,09 € | |||
| Sans rang | Claude InstantPROVISOIREAnthropic | 31,223,0 à 39,4 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Llama 3-8BPROVISOIREMeta AI | 27,819,4 à 36,2 | 33 % · 3 mesures | Inconnu | |||
| Sans rang | Qwen3.5-2BPROVISOIREAlibaba | 24,924,7 à 25,2 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | granite-4.0-1bPROVISOIREIBM | 21,821,6 à 22,0 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | granite-4.0-350mPROVISOIREIBM | 21,821,6 à 22,0 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | DeepSeek LLM 67BPROVISOIREDeepSeek | 19,719,6 à 19,9 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | QwQ-32B-PreviewPROVISOIREAlibaba | 16,914,1 à 19,6 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Llama 2-70BPROVISOIREMeta AI | 14,58,9 à 20,1 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Mistral 7B v0.3PROVISOIREMistral AI | 10,96,8 à 14,9 | 22 % · 2 mesures | 0,22 € / 0,22 € | |||
| Sans rang | phi-3-mini 3.8BPROVISOIREMicrosoft | 10,610,5 à 10,7 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Llama 2-13BPROVISOIREMeta AI | 8,85,2 à 12,4 | 22 % · 2 mesures | Inconnu | |||
| Sans rang | Llama 2-7BPROVISOIREMeta AI | 1,71,7 à 1,7 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | DeepSeek-R1-Distill-Qwen-1.5BPROVISOIREDeepSeek | 0,00,0 à 0,0 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Gemma 3 1BPROVISOIREGoogle DeepMind | 0,00,0 à 0,0 | 11 % · 1 mesure | Inconnu | |||
| Sans rang | Llama 3.2 1BPROVISOIREMeta AI | 0,00,0 à 0,0 | 11 % · 1 mesure | 0,02 € / 0,18 € | |||
Aucun modèle ne correspond à ces filtres.
100 modèles classés et 142 provisoires. La colonne Semaine se remplit à partir de la deuxième édition.
Ce qui compose le score
La démarche →Le score Raisonnement agrège 9 benchmarks, chacun pesé ci-dessous. 1 benchmark archivé ne compte plus. La tâche pèse 15 % du score général.
- Saturé11,1 %ARC-AGI-1
- Saturé11,1 %ARC-AGI-2
- Actif11,1 %Chess Puzzles
- Saturé11,1 %DTBench
- Actif11,1 %EnigmaEval
- Actif11,1 %ForecastBench
- Actif11,1 %LMCA
- Actif11,1 %Mystery Game Puzzles
- Actif11,1 %SimpleBench
Voir aussi : classement général, modèles ouverts, modèles locaux, rapport qualité-prix et votre classement, selon vos priorités.