AI Coding Models Leaderboard
Aggregated official benchmark mirrors and real-world coding task scenarios. 100% verified, objective, and zero paid placement.
📌 7 official boards · mirrored 1:1
Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.
Coding Elo· 51 models- #1Claude Opus 5.51820.3
- #2GPT-6 Astra1792
- #3Claude Fable 5.11753
- #4Claude Sonnet 5.51698.8
- #5Claude Opus 51694
Agent Score· 33 models- #1Claude Fable 5.114.06
- #2Claude Opus 5.511.84
- #3GPT-6 Astra10.36
- #4Claude Opus 59.54
- #5GPT-6 Sol8.8
Pass rate%· 12 models- #1Claude Opus 5.557.8%
- #2Claude Fable 5.151.8%
- #3Claude Opus 546.6%
- #4Grok 4.746.3%
- #5GPT-5.6 Sol41.7%
Intelligence Index· 97 models- #1Claude Opus 5.557.6
- #2Claude Sonnet 5.556
- #3Claude Fable 5.153.4
- #4GPT-6 Astra52.7
- #5Claude Opus 550.8
Score%· 43 models- #1Claude Fable 5.183.4%
- #2Claude Opus 5.583.2%
- #3Claude Fable 583.0%
- #4GPT-6 Astra82.2%
- #5Muse Spark 1.381.6%
Arena Elo· 87 models- #1Claude Opus 5.51508.6
- #2Claude Opus 4.81505.4
- #3Claude Fable 51503.8
- #4Claude Fable 5.11501.3
- #5Muse Spark 1.31496.2
LMArena CodeCoding
Not sure which model fits your budget & stack?
Try 4-Way Plan Showdowns & $20/mo vs API Cost Break-Even Meter
| # | Model | Coding Elo | Value Score | Actions |
|---|---|---|---|---|
| #1 | Claude Opus 5.5Anthropic | 1820.3 | 20 | |
| #2 | GPT-6 AstraOpenAI | 1792 | 8.1 | |
| #3 | Claude Fable 5.1Anthropic | 1753 | 7.8 | |
| #4 | Claude Sonnet 5.5Anthropic | 1698.8 | 25.9 | |
| #5 | Claude Opus 5Anthropic | 1694 | 4.8 | |
| #6 | GPT-6 SolOpenAI | 1692.3 | 37.6 | |
| #7 | Qwen3.8-MaxQwen | 1671 | 47.2 | |
| #8 | Kimi K3Moonshot AI | 1658.5 | 23.8 | |
| #9 | Muse Spark 1.3Meta | 1654.7 | 65.2 | |
| #10 | Grok 4.7xAI | 1636.3 | 43.4 | |
| #11 | Hunyuan-T1Tencent | 1633 | 513.1 | |
| #12 | Claude Fable 5Anthropic | 1625.9 | 7.6 | |
| #13 | GLM-5.3Zhipu AI | 1622.9 | 64.4 | |
| #14 | DeepSeek-V4-FlashDeepSeek | 1620.3 | 418 | |
| #15 | Grok 4.6xAI | 1620.2 | 42 | |
| #16 | GPT-5.6 SolOpenAI | 1619.3 | 11.9 | |
| #17 | GLM-5.2Zhipu AI | 1604.3 | 60.8 | |
| #18 | Gemini 3.7 FlashGoogle | 1593 | 87 | |
| #19 | Qwen3.8-27BQwen | 1591.8 | 271 | |
| #20 | GPT-6 LunaOpenAI | 1583.3 | 633.2 | |
| #21 | DeepSeek-V4-ProDeepSeek | 1582.2 | 130.6 | |
| #22 | Gemini 3.8 FlashGoogle | 1581.4 | 83.5 | |
| #23 | Claude Opus 4.8Anthropic | 1557.1 | 14 | |
| #24 | Grok 4.5xAI | 1551.9 | 42.8 | |
| #25 | Claude Sonnet 5Anthropic | 1539.9 | 30 | |
| #26 | Gemini 3.6 FlashGoogle | 1536.3 | 76.9 | |
| #27 | Claude Sonnet 4.6Anthropic | 1521.4 | 20.9 | |
| #28 | GPT-5.6 TerraOpenAI | 1519.7 | 26.4 | |
| #29 | Doubao Seed 2.0 ProByteDance | 1517.3 | 121.9 | |
| #30 | GPT-5.6 LunaOpenAI | 1516.9 | 246.4 | |
| #31 | Qwen3.7-MaxQwen | 1514.9 | 216.5 | |
| #32 | GPT-5.5OpenAI | 1512.4 | 18.1 | |
| #33 | Kimi K2.6Moonshot AI | 1508.7 | 74.8 | |
| #34 | GLM-5.1Zhipu AI | 1508.6 | 54.8 | |
| #35 | Gemini 3.5 FlashGoogle | 1499.1 | 37.9 | |
| #36 | MiniMax-M3MiniMax | 1482.4 | 199.8 | |
| #37 | Qwen3.6-PlusQwen | 1481.6 | 226.2 | |
| #38 | Kimi K2.7 CodeMoonshot AI | 1473 | 47.6 | |
| #39 | GPT-5.4OpenAI | 1465 | 28.5 | |
| #40 | Gemini 3.1 ProGoogle | 1446.3 | 50.1 | |
| #41 | Gemini 3.5 Flash-LiteGoogle | 1439.7 | 128 | |
| #42 | Gemini 3 ProGoogle | 1439.1 | 44.1 | |
| #43 | GLM-4-PlusZhipu AI | 1434.6 | 74.9 | |
| #44 | GPT-5.4 miniOpenAI | 1397.1 | 417.4 | |
| #45 | MiniMax-M2.7MiniMax | 1397 | 189.5 | |
| #46 | DeepSeek-V3DeepSeek | 1361.7 | 550.6 | |
| #47 | Grok 4.3xAI | 1356.5 | 66.5 | |
| #48 | Claude Haiku 4.5Anthropic | 1329.5 | 43.9 | |
| #49 | Mistral Medium 3.5Mistral AI | 1263.2 | 83.9 | |
| #50 | Mistral Large 3Mistral AI | 1229.9 | 23.9 | |
| #51 | Gemini 2.5 ProGoogle | 1226.9 | 23.9 |
The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.
LMArena Code board: LMSYS Coding Arena Elo ratings based on real developer blind tests on challenging code prompts. Sourced directly from arena.ai/leaderboard/code and mirrored 1:1.
📊 6 benchmark providers
LMArena Code board: LMSYS Coding Arena Elo ratings based on real developer blind tests on challenging code prompts. Sourced directly from arena.ai/leaderboard/code and mirrored 1:1.
Coding Elo51 modelsLMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1.
Agent Score33 modelsCursorBench is Cursor's first-party agentic coding benchmark (v4.0). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.
Pass rate%12 modelsArtificial Analysis Intelligence Index: Tracks reasoning, quality, and coding intelligence across 600+ models, continuously updated as new frontier models release.
Intelligence Index97 modelsLiveBench Coding: Monthly refreshed contamination-resistant suite extracting code completion, generation, JavaScript, Python, and TypeScript performance.
Score%43 modelsLMArena Text board: The global standard LMSYS Chatbot Arena tracking cross-domain text generation and complex instruction-following Elo ratings from arena.ai/leaderboard/text.
Arena Elo87 modelsAll rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.