On Oct 8 JetBrains released Mellum2.1 Thinking: same 12B MoE / 2.5B active / 131K / Apache 2.0 architecture, with gains almost entirely from large-scale RL in real sandboxes. Self-reported same-pipeline scores: SWE-bench Verified 47.0 (vs Mellum2 2.0), LiveCodeBench v6 82.0. On Hugging Face; vLLM serve now; GGUF/MTP coming soon.

Key Takeaways

  • ✓Spec: 12B MoE / 2.5B active, 131K context, Apache 2.0; architecture unchanged vs Mellum2 (HF card)
  • ✓Agentic jump (same pipeline): SWE-bench Verified 2.0→47.0, SWE-bench Pro 0.0→28.0, Terminal-Bench 2.1 0.6→17.4 (Pi v0.73.1)
  • ✓Coding: LiveCodeBench v6 82.0 (leads Qwen3.5-9B 75.4), HumanEval+ 91.5, BFCL v4 62.3
  • ✓Speed: fastest under load in JetBrains’ group; MTP ~1.6× single-request; still trails Qwen3.5-9B on SWE Verified 50.0 / Pro 38.0
  • ✓Ship: vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking --reasoning-parser qwen3; GGUF/Ollama/LM Studio/MTP head coming (blog)
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background

Self-hosted coding agents need small, fast open weights that can actually work inside a repo. JetBrains’ June Mellum2 (12B MoE / 2.5B active) was fast but weak at explore→edit→verify. On Oct 8, Mellum2.1 kept the architecture and poured post-training into large-scale RL across millions of real sandboxes.

Architecture

Unchanged vs Mellum2: 12B MoE / 2.5B active, 131K context, Apache 2.0 (HF card). Gains come from RL scale, harder-filtered task mixes, and SWE training inside real repos with shell/file tools rewarded on passing tests. Positioned as a fast worker/sub-agent for private deploy.

Benchmarks

JetBrains same-pipeline self-reported thinking scores (agentic via Pi v0.73.1): SWE-bench Verified 2.0→47.0, SWE-bench Pro 0.0→28.0, Terminal-Bench 2.1 0.6→17.4. Leads the compared group on LiveCodeBench v6 (82.0), HumanEval+ 91.5, MBPP+ 79.4, BFCL v4 62.3; still trails Qwen3.5-9B on SWE Verified 50.0 / Pro 38.0. MTP ~1.6× single-request; HarmBench 21.5→8.5 (lower better).

Developer path

Serve with vLLM + qwen3 reasoning parser (optional Hermes tool calling). GGUF/Ollama/LM Studio and MTP head are coming soon. Re-benchmark on your own agent harness before production.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.