June 2026

AI Coder Bench

Real-world coding benchmark rankings. 50 tasks across bug fixing, feature building, refactoring, system design, and debug & explain.

7Boards
83Models ranked
777Tasks
August 2026Period
2026-08-15Last update
We fetch raw rankings from 6 public benchmark providers directly — no re-weighting, no composite score. Click "Open official board" on each card to go to the source.

📌 7 official boards · mirrored 1:1

Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider — we do no aggregation or math.

LMArena Agent
Human preference on real agentic coding sessions
Agent
Metric:Agent Score· 30 models
  1. #1Claude Opus 512.19
  2. #2Claude Fable 512.01
  3. #3GPT-5.6 Sol10.86
  4. #4Kimi K310.6
  5. #5Claude Opus 4.89.78
Official ↗
CursorBench
Cursor's first-party agentic coding benchmark (v3.2)
Agent
Metric:Pass rate%· 23 models
  1. #1Claude Fable 572.9%
  2. #2Grok 4.669.9%
  3. #3Gemini 3.7 Flash68.4%
  4. #4GPT-5.6 Sol67.2%
  5. #5Grok 4.566.7%
Official ↗
Artificial Analysis
Composite Intelligence Index over 20+ benchmarks
Overall
Metric:AA Index· 77 models
  1. #1Claude Opus 563.1
  2. #2Claude Fable 562.1
  3. #3GPT-5.6 Sol60.9
  4. #4Grok 4.660.9
  5. #5Kimi K359.7
Official ↗
LiveBench
Contamination-resistant benchmark, refreshed monthly
Reasoning
Metric:Avg score%· 29 models
  1. #1Claude Fable 583.4%
  2. #2GPT-5.6 Sol81.6%
  3. #3GPT-5.580.8%
  4. #4Claude Opus 580.5%
  5. #5Gemini 3.7 Flash79.9%
Official ↗
Aider Polyglot
133 real coding tasks across 6 languages
Coding
Metric:Pass rate%· 35 models
  1. #1GPT-5.6 Sol88.0%
  2. #2o3-pro84.9%
  3. #3Gemini 2.5 Pro83.1%
  4. #4o381.3%
  5. #5Grok 4.679.6%
Official ↗
LiveCodeBench
Continuously updated competitive-programming set
Reasoning
Metric:Pass@1%· 15 models
  1. #1DeepSeek-V3100.0%
  2. #2GPT-4o100.0%
  3. #3GPT-4o mini100.0%
  4. #4o3-mini100.0%
  5. #5o4-mini100.0%
Official ↗
EvalPlus
HumanEval+ strict test suite
Coding
Metric:HumanEval+%· 21 models
  1. #1o189.0%
  2. #2o1-mini89.0%
  3. #3Qwen3.6-Plus87.2%
  4. #4GPT-4o87.2%
  5. #5DeepSeek-V386.6%
Official ↗

LMArena AgentAgent

Metric: Agent Score·30 models·fetched 2026-08-15

Open official board ↗
#ModelAgent ScoreCross-board ranks
#1
Claude Opus 5Anthropic
12.19
C
A #1
L #4
ALE
#212.01
C #1
A #2
L #1
ALE
#310.86
C #4
A #3
L #2
A #1
L
E #21
#4
Kimi K3Moonshot AI
10.6
C #12
A #5
L #6
ALE
#59.78
C #9
A #7
L #13
A #8
L #10
E #11
#6
GPT-5.5OpenAI
8.9
C #17
A #9
L #3
A #19
LE
#77.14
C #13
A #12
L #15
AL
E #8
#8
GLM-5.2Zhipu AI
6.74
C #19
A #15
L #22
ALE
#96.19
C #5
A #11
L #14
ALE
#10
GPT-5.4OpenAI
4.99
C
A #14
L #9
ALE
#114.28
C #14
A #16
L #20
ALE
#124.04
C #15
A #18
L #17
ALE
#133.78
C #7
A #8
L #10
ALE
#143.61
C #3
A #10
L #5
ALE
#153.12
C #21
A #20
L #21
A #13
L #9
E #15
#16
Kimi K2.7 CodeMoonshot AI
1.08
C #20
A #25
L #25
ALE
#17
GLM-5.1Zhipu AI
0.58
C
A #29
LALE
#180.2
C
A #22
L #19
A #15
L #13
E
#190.14
C #10
A #13
L #11
ALE
#20-0.29
C #22
A #17
L #16
ALE
#21-0.44
C
A #21
L #12
ALE
#22
Kimi K2.6Moonshot AI
-0.56
C #23
A #24
L #23
A #16
LE
#23-2.36
C
A #19
L #18
ALE
#24
MiniMax-M3MiniMax
-2.53
C
A #23
L #27
ALE
#25-7.03
C
A #41
LALE
#26-8.47
C
A #34
L #29
ALE
#27-9.1
C
A #31
L #26
ALE
#28-10.1
C
A #26
L
A #3
L #11
E #10
#29-11.17
C
A #33
LALE
#30-14.63
C
A #47
LALE
Rank color:#1#2#3#4-5#6-10#11+

The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.

LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.

📊 7 benchmark providers

LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 — we recompute nothing.

Agent Score30 models

CursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.

Pass rate%23 models

Artificial Analysis aggregates 20+ benchmarks (MMLU-Pro, GPQA, HLE, SciCode, IFBench, Terminal-bench, etc.) into a single Intelligence Index — the most comprehensive commercial evaluator for "overall capability".

AA Index77 models

LiveBench refreshes its questions monthly to prevent contamination. Covers reasoning, math, coding, language, instruction following, and data analysis.

Avg score%29 models

133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.

Pass rate%35 models

Continuously updated competitive-programming problems. Metric: Pass@1.

Pass@1%15 models

HumanEval+ strict test suite, 164 tasks.

HumanEval+%21 models

All rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.