# WeirdML (v2)

> Fiche du benchmark WeirdML (v2) sur Quelle IA, édition du 28 septembre 2026. Écrire du code PyTorch pour des problèmes d'apprentissage automatique inhabituels : reconnaître des formes, classer des chiffres, prédire l'issue de parties d'échecs. En tête au 28 septembre 2026 : GPT-6 Astra (93,6 %). Notation, limites et classement des 129 modèles mesurés.

Page : https://quelleia.com/benchmarks/weirdml/
Source du benchmark : https://htihle.github.io/weirdml.html
Mainteneur : Håvard Tveit Ihle
Tâche : Code
État : actif

## À quoi il sert

Sur Quelle IA, WeirdML (v2) sert à noter la tâche Code : c'est l'un de ses 13 benchmarks. Il départage surtout les modèles dont le score IA approche 75.

## Comment c'est noté

Précision moyenne obtenue par le code du modèle, exécuté dans un environnement isolé, avec cinq tentatives par tâche.

## Ce qu'il ne dit pas

Version remplacée en septembre 2026 par un nouveau jeu de tâches. Ses scores ne se comparent pas à ceux de WeirdML v3.

## Qui le tient

Håvard Tveit Ihle.

## Comment lire le score

Chaque trait est un modèle, placé à son meilleur score. Un modèle qui obtient 50 % a un score IA d'environ 75. Le meilleur, GPT-6 Astra, atteint 93,6 %. Le benchmark sera dit saturé quand un modèle atteindra 95 %.

Ce que le benchmark attend à chaque niveau, d'après sa courbe d'étalonnage :

- Score IA 51 (niveau de GPT-4o (Nov 2024)) : 20 %
- Score IA 76 (niveau de Claude Sonnet 4.5) : 53 %
- Score IA 100 (niveau de Claude Opus 5.5) : 82 %

## Son poids dans le score IA

1,15 % du score général. WeirdML (v2) porte 8 % de la tâche Code (15 % du score général), que se partagent 13 benchmarks.

## Classement (129 modèles mesurés · 172 mesures, édition du 28 septembre 2026)

Un modèle par ligne, à sa meilleure configuration. « Suggère » : le score IA que cette seule mesure indique.

| Rang | Modèle | Éditeur | Réflexion | Score publié | Suggère | Source |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-6 Astra](https://quelleia.com/modeles/gpt-6-astra.md) | OpenAI | Maximale | 93,6 % | 119,9 | https://htihle.github.io/weirdml.html |
| 2 | [Claude Fable 5.1](https://quelleia.com/modeles/claude-fable-5-1.md) | Anthropic | Maximale | 92,9 % | 118,1 | https://htihle.github.io/weirdml.html |
| 3 | [Claude Fable 5](https://quelleia.com/modeles/claude-fable-5.md) | Anthropic | Maximale | 91,9 % | 115,8 | https://htihle.github.io/weirdml.html |
| 4 | [Claude Opus 5](https://quelleia.com/modeles/claude-opus-5.md) | Anthropic | Maximale | 91,8 % | 115,4 | https://htihle.github.io/weirdml.html |
| 5 | [GPT-5.6 Sol](https://quelleia.com/modeles/gpt-5-6-sol.md) | OpenAI | Maximale | 89,4 % | 110,7 | https://htihle.github.io/weirdml.html |
| 6 | [GPT-5.5](https://quelleia.com/modeles/gpt-5-5.md) | OpenAI | Maximale | 84,9 % | 103,8 | https://htihle.github.io/weirdml.html |
| 7 | [Claude Opus 4.8](https://quelleia.com/modeles/claude-opus-4-8.md) | Anthropic | Maximale | 82,9 % | 101,3 | https://htihle.github.io/weirdml.html |
| 8 | [Kimi K3](https://quelleia.com/modeles/kimi-k3.md) | Moonshot | Maximale | 82,6 % | 100,9 | https://htihle.github.io/weirdml.html |
| 9 | [GPT-5.3 Codex](https://quelleia.com/modeles/gpt-5-3-codex.md) | OpenAI | non précisée | 79,3 % | 97,3 | https://epoch.ai/benchmarks |
| 10 | [GPT-5.6 Terra](https://quelleia.com/modeles/gpt-5-6-terra.md) | OpenAI | Élevée | 78,3 % | 96,3 | https://htihle.github.io/weirdml.html |
| 11 | [Claude Opus 4.6](https://quelleia.com/modeles/claude-opus-4-6.md) | Anthropic | Élevée | 78 % | 95,9 | https://htihle.github.io/weirdml.html |
| 12 | [GPT-5.4](https://quelleia.com/modeles/gpt-5-4.md) | OpenAI | Maximale | 77,7 % | 95,7 | https://htihle.github.io/weirdml.html |
| 13 | [Claude Opus 4.7](https://quelleia.com/modeles/claude-opus-4-7.md) | Anthropic | Élevée | 76,4 % | 94,5 | https://htihle.github.io/weirdml.html |
| 14 | [GLM-5.3](https://quelleia.com/modeles/glm-5-3.md) | Z.ai (Zhipu AI) | Maximale | 75,4 % | 93,5 | https://htihle.github.io/weirdml.html |
| 15 | [GPT-5.2](https://quelleia.com/modeles/gpt-5-2.md) | OpenAI | Maximale | 72,2 % | 90,7 | https://htihle.github.io/weirdml.html |
| 16 | [Gemini 3.1 Pro](https://quelleia.com/modeles/gemini-3-1-pro.md) | Google DeepMind | non précisée | 72,1 % | 90,6 | https://htihle.github.io/weirdml.html |
| 17 | [GLM-5.2](https://quelleia.com/modeles/glm-5-2.md) | Z.ai (Zhipu AI) | Maximale | 70,1 % | 89,0 | https://htihle.github.io/weirdml.html |
| 18 | [Gemini 3 Pro](https://quelleia.com/modeles/gemini-3-pro.md) | Google DeepMind | non précisée | 69,9 % | 88,9 | https://htihle.github.io/weirdml.html |
| 19 | [Claude Sonnet 5](https://quelleia.com/modeles/claude-sonnet-5.md) | Anthropic | Élevée | 68,8 % | 87,9 | https://htihle.github.io/weirdml.html |
| 20 | [Grok 4.6](https://quelleia.com/modeles/grok-4-6.md) | xAI | Élevée | 67,3 % | 86,8 | https://htihle.github.io/weirdml.html |
| 21 | [DeepSeek V4 Pro 0813](https://quelleia.com/modeles/deepseek-v4-pro-0813.md) | DeepSeek | Maximale | 66,2 % | 86,0 | https://htihle.github.io/weirdml.html |
| 22 | [Claude Sonnet 4.6](https://quelleia.com/modeles/claude-sonnet-4-6.md) | Anthropic | Moyenne | 66,1 % | 85,8 | https://epoch.ai/benchmarks |
| 23 | [Claude Opus 4.5](https://quelleia.com/modeles/claude-opus-4-5.md) | Anthropic | Budget | 63,7 % | 84,1 | https://htihle.github.io/weirdml.html |
| 24 | [DeepSeek V4 Flash 0731](https://quelleia.com/modeles/deepseek-v4-flash-0731.md) | DeepSeek | Maximale | 63 % | 83,6 | https://htihle.github.io/weirdml.html |
| 25 | [Gemini 3.5 Flash](https://quelleia.com/modeles/gemini-3-5-flash.md) | Google DeepMind | Élevée | 62,6 % | 83,3 | https://htihle.github.io/weirdml.html |
| 26 | [Gemini 3 Flash](https://quelleia.com/modeles/gemini-3-flash.md) | Google DeepMind | non précisée | 61,6 % | 82,6 | https://htihle.github.io/weirdml.html |
| 27 | [GPT-5.6 Luna](https://quelleia.com/modeles/gpt-5-6-luna.md) | OpenAI | Élevée | 60,9 % | 82,0 | https://htihle.github.io/weirdml.html |
| 28 | [GPT-5.1](https://quelleia.com/modeles/gpt-5-1.md) | OpenAI | Élevée | 60,8 % | 82,0 | https://htihle.github.io/weirdml.html |
| 29 | [GPT-5](https://quelleia.com/modeles/gpt-5.md) | OpenAI | Élevée | 60,7 % | 81,9 | https://htihle.github.io/weirdml.html |
| 30 | [GPT-5 Pro](https://quelleia.com/modeles/gpt-5-pro.md) | OpenAI | Élevée | 60,4 % | 81,7 | https://htihle.github.io/weirdml.html |
| 31 | [Muse Spark 1.2](https://quelleia.com/modeles/muse-spark-1-2.md) | Meta AI | Maximale | 60,3 % | 81,6 | https://htihle.github.io/weirdml.html |
| 31 | [GPT-5.4 Mini](https://quelleia.com/modeles/gpt-5-4-mini.md) | OpenAI | Élevée | 60,3 % | 81,6 | https://htihle.github.io/weirdml.html |
| 33 | [o3-pro](https://quelleia.com/modeles/o3-pro.md) | OpenAI | Élevée | 58,2 % | 80,2 | https://htihle.github.io/weirdml.html |
| 34 | [GPT-5.4 Pro](https://quelleia.com/modeles/gpt-5-4-pro.md) | OpenAI | Sans | 57,4 % | 79,6 | https://htihle.github.io/weirdml.html |
| 35 | [GLM-5.1](https://quelleia.com/modeles/glm-5-1.md) | Z.ai (Zhipu AI) | non précisée | 57,1 % | 79,4 | https://htihle.github.io/weirdml.html |
| 36 | [Gemini 3.6 Flash](https://quelleia.com/modeles/gemini-3-6-flash.md) | Google DeepMind | Élevée | 56,1 % | 78,7 | https://htihle.github.io/weirdml.html |
| 37 | [Kimi K2.6](https://quelleia.com/modeles/kimi-k2-6.md) | Moonshot | non précisée | 55,9 % | 78,6 | https://htihle.github.io/weirdml.html |
| 38 | [GPT-5-Codex](https://quelleia.com/modeles/gpt-5-codex.md) | OpenAI | non précisée | 54,5 % | 77,7 | https://epoch.ai/benchmarks |
| 39 | [Kimi K2.7 Code](https://quelleia.com/modeles/kimi-k2-7-code.md) | Moonshot | non précisée | 54,1 % | 77,4 | https://htihle.github.io/weirdml.html |
| 40 | [Gemini 2.5 Pro (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-pro-jun-2025.md) | Google DeepMind | Budget | 54 % | 77,3 | https://epoch.ai/benchmarks |
| 41 | [GPT-5 mini](https://quelleia.com/modeles/gpt-5-mini.md) | OpenAI | Élevée | 52,7 % | 76,4 | https://htihle.github.io/weirdml.html |
| 42 | [o4-mini](https://quelleia.com/modeles/o4-mini.md) | OpenAI | Élevée | 52,6 % | 76,3 | https://htihle.github.io/weirdml.html |
| 43 | [o3](https://quelleia.com/modeles/o3.md) | OpenAI | Élevée | 52,4 % | 76,2 | https://htihle.github.io/weirdml.html |
| 44 | [Grok 4.20](https://quelleia.com/modeles/grok-4-20.md) | xAI | non précisée | 52,3 % | 76,1 | https://htihle.github.io/weirdml.html |
| 44 | [Gemma 4 31B IT](https://quelleia.com/modeles/gemma-4-31b-it.md) | Google DeepMind | non précisée | 52,3 % | 76,1 | https://htihle.github.io/weirdml.html |
| 46 | [Gemini 3.1 Flash-Lite](https://quelleia.com/modeles/gemini-3-1-flash-lite.md) | Google DeepMind | non précisée | 52,2 % | 76,1 | https://epoch.ai/benchmarks |
| 47 | [Grok 4.3 Beta](https://quelleia.com/modeles/grok-4-3-beta.md) | xAI | non précisée | 49,9 % | 74,5 | https://htihle.github.io/weirdml.html |
| 48 | [GPT-5.4 Nano](https://quelleia.com/modeles/gpt-5-4-nano.md) | OpenAI | Élevée | 49,2 % | 74,1 | https://htihle.github.io/weirdml.html |
| 49 | [DeepSeek-V4-Pro](https://quelleia.com/modeles/deepseek-v4-pro.md) | DeepSeek | Maximale | 48,9 % | 73,8 | https://htihle.github.io/weirdml.html |
| 50 | [GLM-5](https://quelleia.com/modeles/glm-5.md) | Z.ai (Zhipu AI) | non précisée | 48,2 % | 73,3 | https://epoch.ai/benchmarks |
| 50 | [gpt-oss-120b](https://quelleia.com/modeles/gpt-oss-120b.md) | OpenAI | Élevée | 48,2 % | 73,3 | https://epoch.ai/benchmarks |
| 52 | [Claude Sonnet 4.5](https://quelleia.com/modeles/claude-sonnet-4-5.md) | Anthropic | Budget | 47,7 % | 73,0 | https://htihle.github.io/weirdml.html |
| 53 | [o1-preview](https://quelleia.com/modeles/o1-preview.md) | OpenAI | non précisée | 47,6 % | 72,9 | https://htihle.github.io/weirdml.html |
| 54 | [DeepSeek-V3.2-Speciale](https://quelleia.com/modeles/deepseek-v3-2-speciale.md) | DeepSeek | non précisée | 46,7 % | 72,4 | https://epoch.ai/benchmarks |
| 55 | [Grok 4.5](https://quelleia.com/modeles/grok-4-5.md) | xAI | non précisée | 46,4 % | 72,2 | https://htihle.github.io/weirdml.html |
| 56 | [Claude Sonnet 4](https://quelleia.com/modeles/claude-sonnet-4.md) | Anthropic | Budget | 46,1 % | 71,9 | https://htihle.github.io/weirdml.html |
| 57 | [o1](https://quelleia.com/modeles/o1.md) | OpenAI | Élevée | 46,1 % | 71,9 | https://htihle.github.io/weirdml.html |
| 58 | [Claude Opus 4.1](https://quelleia.com/modeles/claude-opus-4-1.md) | Anthropic | Budget | 45,9 % | 71,8 | https://htihle.github.io/weirdml.html |
| 59 | [Grok 4](https://quelleia.com/modeles/grok-4.md) | xAI | non précisée | 45,7 % | 71,7 | https://htihle.github.io/weirdml.html |
| 60 | [DeepSeek-V4-Flash](https://quelleia.com/modeles/deepseek-v4-flash.md) | DeepSeek | Maximale | 45,6 % | 71,6 | https://htihle.github.io/weirdml.html |
| 61 | [Kimi K2.5](https://quelleia.com/modeles/kimi-k2-5.md) | Moonshot | non précisée | 45,6 % | 71,6 | https://htihle.github.io/weirdml.html |
| 62 | [Claude Haiku 4.5](https://quelleia.com/modeles/claude-haiku-4-5.md) | Anthropic | non précisée | 45,4 % | 71,5 | https://htihle.github.io/weirdml.html |
| 63 | [Claude Opus 4](https://quelleia.com/modeles/claude-opus-4.md) | Anthropic | Budget | 43,7 % | 70,3 | https://htihle.github.io/weirdml.html |
| 64 | [Mistral Medium 3.5](https://quelleia.com/modeles/mistral-medium-3-5.md) | Mistral AI | non précisée | 43,7 % | 70,3 | https://htihle.github.io/weirdml.html |
| 65 | [o3-mini](https://quelleia.com/modeles/o3-mini.md) | OpenAI | Élevée | 43,7 % | 70,3 | https://htihle.github.io/weirdml.html |
| 66 | [Nemotron 3 Ultra](https://quelleia.com/modeles/nemotron-3-ultra.md) | NVIDIA | non précisée | 43,5 % | 70,1 | https://htihle.github.io/weirdml.html |
| 67 | [Mercury 2](https://quelleia.com/modeles/mercury-2.md) | Inception Labs | non précisée | 43,2 % | 69,9 | https://htihle.github.io/weirdml.html |
| 68 | [Grok 4 Fast](https://quelleia.com/modeles/grok-4-fast.md) | xAI | non précisée | 42,9 % | 69,7 | https://htihle.github.io/weirdml.html |
| 69 | [Kimi K2 Thinking](https://quelleia.com/modeles/kimi-k2-thinking.md) | Moonshot | non précisée | 42,8 % | 69,7 | https://htihle.github.io/weirdml.html |
| 70 | [Grok-3 mini](https://quelleia.com/modeles/grok-3-mini.md) | xAI | Élevée | 42,6 % | 69,5 | https://epoch.ai/benchmarks |
| 71 | [Gemini 2.5 Flash (Sep 2025)](https://quelleia.com/modeles/gemini-2-5-flash-sep-2025.md) | Google DeepMind | Budget | 41,9 % | 69,1 | https://htihle.github.io/weirdml.html |
| 72 | [DeepSeek-R1 (May 2025)](https://quelleia.com/modeles/deepseek-r1-may-2025.md) | DeepSeek | non précisée | 41,6 % | 68,9 | https://htihle.github.io/weirdml.html |
| 73 | [Qwen3-Coder-480B-A35B](https://quelleia.com/modeles/qwen3-coder-480b-a35b.md) | Alibaba | non précisée | 41,2 % | 68,5 | https://htihle.github.io/weirdml.html |
| 74 | [Qwen3-235B-A22B-Thinking (Jul 2025)](https://quelleia.com/modeles/qwen3-235b-a22b-thinking-jul-2025.md) | Alibaba | non précisée | 41 % | 68,4 | https://epoch.ai/benchmarks |
| 75 | [Gemini 2.5 Flash (Apr 2025)](https://quelleia.com/modeles/gemini-2-5-flash-apr-2025.md) | Google DeepMind | non précisée | 40,9 % | 68,4 | https://epoch.ai/benchmarks |
| 75 | [Gemini 2.5 Flash (May 2025)](https://quelleia.com/modeles/gemini-2-5-flash-may-2025.md) | Google DeepMind | Budget | 40,9 % | 68,4 | https://epoch.ai/benchmarks |
| 77 | [gpt-oss-20b](https://quelleia.com/modeles/gpt-oss-20b.md) | OpenAI | Élevée | 40,9 % | 68,4 | https://htihle.github.io/weirdml.html |
| 78 | [GLM-4.5](https://quelleia.com/modeles/glm-4-5.md) | Z.ai (Zhipu AI) | Élevée | 40,6 % | 68,1 | https://htihle.github.io/weirdml.html |
| 79 | [Claude 3.5 Sonnet (October 2024)](https://quelleia.com/modeles/claude-3-5-sonnet-october-2024.md) | Anthropic | non précisée | 40 % | 67,7 | https://htihle.github.io/weirdml.html |
| 80 | [Qwen3.5-27B](https://quelleia.com/modeles/qwen3-5-27b.md) | Alibaba | non précisée | 39,5 % | 67,4 | https://htihle.github.io/weirdml.html |
| 81 | [DeepSeek-V3.2-Exp](https://quelleia.com/modeles/deepseek-v3-2-exp.md) | DeepSeek | Élevée | 39,5 % | 67,4 | https://htihle.github.io/weirdml.html |
| 82 | [GPT-4.5](https://quelleia.com/modeles/gpt-4-5.md) | OpenAI | non précisée | 39,4 % | 67,3 | https://htihle.github.io/weirdml.html |
| 83 | [Kimi K2 (Jul 2025)](https://quelleia.com/modeles/kimi-k2-jul-2025.md) | Moonshot | non précisée | 39,4 % | 67,3 | https://epoch.ai/benchmarks |
| 84 | [GPT-4.1](https://quelleia.com/modeles/gpt-4-1.md) | OpenAI | non précisée | 39 % | 67,0 | https://htihle.github.io/weirdml.html |
| 85 | [Gemini 3.5 Flash-Lite](https://quelleia.com/modeles/gemini-3-5-flash-lite.md) | Google DeepMind | Élevée | 39 % | 67,0 | https://htihle.github.io/weirdml.html |
| 86 | [Qwen3-235B-A22B-Instruct (Jul 2025)](https://quelleia.com/modeles/qwen3-235b-a22b-instruct-jul-2025.md) | Alibaba | non précisée | 38,7 % | 66,8 | https://epoch.ai/benchmarks |
| 87 | [DeepSeek-V3.1](https://quelleia.com/modeles/deepseek-v3-1.md) | DeepSeek | Élevée | 38,4 % | 66,6 | https://htihle.github.io/weirdml.html |
| 88 | [GPT-5 nano](https://quelleia.com/modeles/gpt-5-nano.md) | OpenAI | Élevée | 38,1 % | 66,3 | https://htihle.github.io/weirdml.html |
| 89 | [Nemotron 3 Super](https://quelleia.com/modeles/nemotron-3-super.md) | NVIDIA | non précisée | 38 % | 66,3 | https://htihle.github.io/weirdml.html |
| 90 | [GPT-4.1 mini](https://quelleia.com/modeles/gpt-4-1-mini.md) | OpenAI | non précisée | 37,6 % | 66,0 | https://htihle.github.io/weirdml.html |
| 91 | [Qwen3-235B-A22B](https://quelleia.com/modeles/qwen3-235b-a22b.md) | Alibaba | non précisée | 37,3 % | 65,8 | https://epoch.ai/benchmarks |
| 92 | [Grok 3](https://quelleia.com/modeles/grok-3.md) | xAI | non précisée | 37,2 % | 65,7 | https://htihle.github.io/weirdml.html |
| 93 | [MiniMax-M2.7](https://quelleia.com/modeles/minimax-m2-7.md) | MiniMax | non précisée | 37 % | 65,5 | https://htihle.github.io/weirdml.html |
| 94 | [Kimi K2 (Sep 2025)](https://quelleia.com/modeles/kimi-k2-sep-2025.md) | Moonshot | non précisée | 36,7 % | 65,3 | https://htihle.github.io/weirdml.html |
| 95 | [DeepSeek-R1](https://quelleia.com/modeles/deepseek-r1.md) | DeepSeek | non précisée | 36,5 % | 65,2 | https://htihle.github.io/weirdml.html |
| 96 | [o1-mini](https://quelleia.com/modeles/o1-mini.md) | OpenAI | Moyenne | 36,3 % | 65,1 | https://htihle.github.io/weirdml.html |
| 97 | [DeepSeek-V3 (Mar 2025)](https://quelleia.com/modeles/deepseek-v3-mar-2025.md) | DeepSeek | non précisée | 36,1 % | 64,9 | https://htihle.github.io/weirdml.html |
| 98 | [Gemini 2.5 Flash-Lite (Jun 2025)](https://quelleia.com/modeles/gemini-2-5-flash-lite-jun-2025.md) | Google DeepMind | Budget | 35,2 % | 64,3 | https://epoch.ai/benchmarks |
| 99 | [Gemma 4 26B A4B](https://quelleia.com/modeles/gemma-4-26b-a4b.md) | Google DeepMind | non précisée | 35,2 % | 64,2 | https://htihle.github.io/weirdml.html |
| 100 | [grok-code-fast-1](https://quelleia.com/modeles/grok-code-fast-1.md) | xAI | non précisée | 35,1 % | 64,2 | https://htihle.github.io/weirdml.html |
| 101 | [Qwen 3.6 35B-A3B](https://quelleia.com/modeles/qwen-3-6-35b-a3b.md) | Alibaba | non précisée | 34,5 % | 63,7 | https://htihle.github.io/weirdml.html |
| 102 | [qwen3-coder-next](https://quelleia.com/modeles/qwen3-coder-next.md) | Alibaba | non précisée | 34,4 % | 63,7 | https://epoch.ai/benchmarks |
| 103 | [Mistral Medium 3.1](https://quelleia.com/modeles/mistral-medium-3-1.md) | Mistral AI | non précisée | 33,1 % | 62,7 | https://htihle.github.io/weirdml.html |
| 104 | [Inkling](https://quelleia.com/modeles/inkling.md) | Thinking Machines | Élevée | 32,3 % | 62,1 | https://htihle.github.io/weirdml.html |
| 105 | [Claude 3.5 Sonnet](https://quelleia.com/modeles/claude-3-5-sonnet.md) | Anthropic | non précisée | 31 % | 61,0 | https://htihle.github.io/weirdml.html |
| 106 | [Claude 3.5 Haiku](https://quelleia.com/modeles/claude-3-5-haiku.md) | Anthropic | non précisée | 30,7 % | 60,8 | https://htihle.github.io/weirdml.html |
| 107 | [Qwen3-30B-A3B](https://quelleia.com/modeles/qwen3-30b-a3b.md) | Alibaba | non précisée | 29,8 % | 60,0 | https://epoch.ai/benchmarks |
| 108 | [Gemini 2.0 Flash (Feb 2025)](https://quelleia.com/modeles/gemini-2-0-flash-feb-2025.md) | Google DeepMind | non précisée | 25,8 % | 56,7 | https://htihle.github.io/weirdml.html |
| 109 | [GPT-4o (Nov 2024)](https://quelleia.com/modeles/gpt-4o-nov-2024.md) | OpenAI | non précisée | 25,1 % | 56,1 | https://htihle.github.io/weirdml.html |
| 110 | [Gemini 1.5 Flash (Sep 2024)](https://quelleia.com/modeles/gemini-1-5-flash-sep-2024.md) | Google DeepMind | non précisée | 24,9 % | 55,9 | https://htihle.github.io/weirdml.html |
| 111 | [Llama 4 Maverick](https://quelleia.com/modeles/llama-4-maverick.md) | Meta AI | non précisée | 24,5 % | 55,5 | https://htihle.github.io/weirdml.html |
| 112 | [Grok-2 (Dec 2024)](https://quelleia.com/modeles/grok-2-dec-2024.md) | xAI | non précisée | 22,2 % | 53,4 | https://htihle.github.io/weirdml.html |
| 113 | [Gemini 1.5 Pro (Sept 2024)](https://quelleia.com/modeles/gemini-1-5-pro-sept-2024.md) | Google DeepMind | non précisée | 22,2 % | 53,4 | https://htihle.github.io/weirdml.html |
| 114 | [Llama 3.1-405B](https://quelleia.com/modeles/llama-3-1-405b.md) | Meta AI | non précisée | 21,4 % | 52,6 | https://htihle.github.io/weirdml.html |
| 115 | [Claude 3 Opus](https://quelleia.com/modeles/claude-3-opus.md) | Anthropic | non précisée | 19,2 % | 50,3 | https://htihle.github.io/weirdml.html |
| 116 | [GPT-4.1 nano](https://quelleia.com/modeles/gpt-4-1-nano.md) | OpenAI | non précisée | 19 % | 50,0 | https://htihle.github.io/weirdml.html |
| 117 | [GPT-4 Turbo (Apr 2024)](https://quelleia.com/modeles/gpt-4-turbo-apr-2024.md) | OpenAI | non précisée | 18 % | 48,9 | https://htihle.github.io/weirdml.html |
| 118 | [Qwen2.5-72B](https://quelleia.com/modeles/qwen2-5-72b.md) | Alibaba | non précisée | 16 % | 46,5 | https://epoch.ai/benchmarks |
| 119 | [Llama 3.3 70B](https://quelleia.com/modeles/llama-3-3-70b.md) | Meta AI | non précisée | 14,4 % | 44,5 | https://htihle.github.io/weirdml.html |
| 120 | [GPT-4 (Jun 2023)](https://quelleia.com/modeles/gpt-4-jun-2023.md) | OpenAI | non précisée | 12,4 % | 41,4 | https://htihle.github.io/weirdml.html |
| 121 | [GPT-4o mini](https://quelleia.com/modeles/gpt-4o-mini.md) | OpenAI | non précisée | 11,8 % | 40,5 | https://htihle.github.io/weirdml.html |
| 122 | [Qwen2-72B](https://quelleia.com/modeles/qwen2-72b.md) | Alibaba | non précisée | 11,3 % | 39,7 | https://htihle.github.io/weirdml.html |
| 123 | [Claude 3 Sonnet](https://quelleia.com/modeles/claude-3-sonnet.md) | Anthropic | non précisée | 10,2 % | 37,7 | https://htihle.github.io/weirdml.html |
| 124 | [Claude 3 Haiku](https://quelleia.com/modeles/claude-3-haiku.md) | Anthropic | non précisée | 9,8 % | 37,1 | https://htihle.github.io/weirdml.html |
| 125 | [Llama 3.1-70B](https://quelleia.com/modeles/llama-3-1-70b.md) | Meta AI | non précisée | 9 % | 35,4 | https://htihle.github.io/weirdml.html |
| 126 | [Claude 2.1](https://quelleia.com/modeles/claude-2-1.md) | Anthropic | non précisée | 7,1 % | 31,0 | https://htihle.github.io/weirdml.html |
| 127 | [GPT-3.5 Turbo (Jan 2024)](https://quelleia.com/modeles/gpt-3-5-turbo-jan-2024.md) | OpenAI | non précisée | 3,5 % | 18,4 | https://htihle.github.io/weirdml.html |
| 128 | [Mixtral 8x22B](https://quelleia.com/modeles/mixtral-8x22b.md) | Mistral AI | non précisée | 3,2 % | 16,7 | https://htihle.github.io/weirdml.html |
| 129 | [Llama 3.1-8B](https://quelleia.com/modeles/llama-3-1-8b.md) | Meta AI | non précisée | 1,7 % | 6,2 | https://htihle.github.io/weirdml.html |

## En bref

**Que mesure WeirdML (v2) ?** Écrire du code PyTorch pour des problèmes d'apprentissage automatique inhabituels : reconnaître des formes, classer des chiffres, prédire l'issue de parties d'échecs. Sur Quelle IA, il est rangé dans la tâche Code.

**Comment WeirdML (v2) est-il noté ?** Précision moyenne obtenue par le code du modèle, exécuté dans un environnement isolé, avec cinq tentatives par tâche. Un modèle qui obtient 50 % a un score IA d'environ 75.

**Qui est en tête sur WeirdML (v2) ?** GPT-6 Astra (OpenAI) est en tête de WeirdML (v2) avec 93,6 %, dans l'édition du 28 septembre 2026. Suivent Claude Fable 5.1 (92,9 %) et Claude Fable 5 (91,9 %). 129 modèles y sont mesurés.

**Quelles sont les limites de WeirdML (v2) ?** Version remplacée en septembre 2026 par un nouveau jeu de tâches. Ses scores ne se comparent pas à ceux de WeirdML v3. Au 28 septembre 2026, le benchmark est actif.

**Combien pèse WeirdML (v2) dans le score IA ?** WeirdML (v2) compte pour 1,15 % du score général, dans l'édition du 28 septembre 2026. Il porte 8 % de la tâche Code (15 % du score général), que se partagent 13 benchmarks.

**Qui maintient WeirdML (v2) ?** Håvard Tveit Ihle. Sa page de référence : htihle.github.io/weirdml.html.

## Les autres benchmarks de la tâche Code

- [Arena, questions de code](https://quelleia.com/benchmarks/arena_code.md) : 275 modèles · 1,15 % du score
- [SciCode](https://quelleia.com/benchmarks/scicode.md) : 120 modèles · 1,15 % du score
- [Terminal-Bench 2.0](https://quelleia.com/benchmarks/terminal_bench.md) : 45 modèles · 1,15 % du score
- [FrontierCode](https://quelleia.com/benchmarks/frontiercode.md) : 37 modèles · 1,15 % du score
- [SWE-bench Verified](https://quelleia.com/benchmarks/swe_bench_verified.md) : 32 modèles · 1,15 % du score
- [DeepSWE](https://quelleia.com/benchmarks/deepswe.md) : 26 modèles · 1,15 % du score
- [GSO](https://quelleia.com/benchmarks/gso_bench.md) : 26 modèles · 1,15 % du score
- [GBAEval](https://quelleia.com/benchmarks/gbaeval.md) : 23 modèles · 1,15 % du score
- [FrontierSWE](https://quelleia.com/benchmarks/frontierswe.md) : 16 modèles · 1,15 % du score
- [CursorBench](https://quelleia.com/benchmarks/cursorbench.md) : 13 modèles · 1,15 % du score
- [WeirdML v3](https://quelleia.com/benchmarks/weirdml_v3.md) : 11 modèles · 1,15 % du score
- [MirrorCode](https://quelleia.com/benchmarks/mirrorcode.md) : 9 modèles · 1,15 % du score
- [Aider Polyglot](https://quelleia.com/benchmarks/aider_polyglot.md) : 54 modèles · hors score, archivé
- [LiveBench, code](https://quelleia.com/benchmarks/livebench_code.md) : 48 modèles · hors score, archivé

Source : Quelle IA, édition du 28 septembre 2026. https://quelleia.com
