AI Coder Bench
Real-world coding benchmark rankings. 50 tasks across bug fixing, feature building, refactoring, system design, and debug & explain.
📌 7 official boards · mirrored 1:1
Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.
Agent Score· 30 models- #1Claude Opus 512.19
- #2Claude Fable 512.01
- #3GPT-5.6 Sol10.86
- #4Kimi K310.6
- #5Claude Opus 4.89.78
Pass rate%· 23 models- #1Claude Fable 572.9%
- #2Grok 4.669.9%
- #3Gemini 3.7 Flash68.4%
- #4GPT-5.6 Sol67.2%
- #5Grok 4.566.7%
AA Index· 77 models- #1Claude Opus 563.1
- #2Claude Fable 562.1
- #3GPT-5.6 Sol60.9
- #4Grok 4.660.9
- #5Kimi K359.7
Avg score%· 29 models- #1Claude Fable 583.4%
- #2GPT-5.6 Sol81.6%
- #3GPT-5.580.8%
- #4Claude Opus 580.5%
- #5Gemini 3.7 Flash79.9%
Pass rate%· 35 models- #1GPT-5.6 Sol88.0%
- #2o3-pro84.9%
- #3Gemini 2.5 Pro83.1%
- #4o381.3%
- #5Grok 4.679.6%
Pass@1%· 15 models- #1DeepSeek-V3100.0%
- #2GPT-4o100.0%
- #3GPT-4o mini100.0%
- #4o3-mini100.0%
- #5o4-mini100.0%
HumanEval+%· 21 models- #1o189.0%
- #2o1-mini89.0%
- #3Qwen3.6-Plus87.2%
- #4GPT-4o87.2%
- #5DeepSeek-V386.6%
LMArena AgentAgent
| # | Model | Agent Score |
|---|---|---|
| #1 | Claude Opus 5Anthropic | 12.19 |
| #2 | Claude Fable 5Anthropic | 12.01 |
| #3 | GPT-5.6 SolOpenAI | 10.86 |
| #4 | Kimi K3Moonshot AI | 10.6 |
| #5 | Claude Opus 4.8Anthropic | 9.78 |
| #6 | GPT-5.5OpenAI | 8.9 |
| #7 | Claude Sonnet 5Anthropic | 7.14 |
| #8 | GLM-5.2Zhipu AI | 6.74 |
| #9 | Grok 4.5xAI | 6.19 |
| #10 | GPT-5.4OpenAI | 4.99 |
| #11 | GPT-5.6 LunaOpenAI | 4.28 |
| #12 | DeepSeek-V4-FlashDeepSeek | 4.04 |
| #13 | GPT-5.6 TerraOpenAI | 3.78 |
| #14 | Gemini 3.7 FlashGoogle | 3.61 |
| #15 | Claude Sonnet 4.6Anthropic | 3.12 |
| #16 | Kimi K2.7 CodeMoonshot AI | 1.08 |
| #17 | GLM-5.1Zhipu AI | 0.58 |
| #18 | Qwen3.7-MaxQwen | 0.2 |
| #19 | DeepSeek-V4-ProDeepSeek | 0.14 |
| #20 | Gemini 3.5 FlashGoogle | -0.29 |
| #21 | Gemini 3.1 ProGoogle | -0.44 |
| #22 | Kimi K2.6Moonshot AI | -0.56 |
| #23 | Gemini 3 ProGoogle | -2.36 |
| #24 | MiniMax-M3MiniMax | -2.53 |
| #25 | Mistral Medium 3.5Mistral AI | -7.03 |
| #26 | Grok 4.3xAI | -8.47 |
| #27 | -9.1 | |
| #28 | Gemini 2.5 ProGoogle | -10.1 |
| #29 | MiniMax-M2.7MiniMax | -11.17 |
| #30 | Nemotron-UltraNvidia | -14.63 |
The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
📊 7 benchmark providers
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.
Agent Score30 modelsCursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.
Pass rate%23 modelsArtificial Analysis aggregates 20+ benchmarks (MMLU-Pro, GPQA, HLE, SciCode, IFBench, Terminal-bench, etc.) into a single Intelligence Index — the most comprehensive commercial evaluator for "overall capability".
AA Index77 modelsLiveBench refreshes its questions monthly to prevent contamination. Covers reasoning, math, coding, language, instruction following, and data analysis.
Avg score%29 models133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.
Pass rate%35 modelsContinuously updated competitive-programming problems. Metric: Pass@1.
Pass@1%15 modelsAll rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.