On Oct 5 Reflection AI introduced Beam, a sparse MoE with 501B total and 23B active parameters for coding, reasoning and agentic work. It was pretrained on 23.8T tokens, then put through a four-week high-compute RL run on 10.5K GB300 GPUs (100M+ rollouts), with 1M-token context. Reflection says it is competitive with GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, matching GLM-5.2 on reasoning with 3–4x less inference compute. It is in final red-teaming with early-access signup only; weights, tech report and model card ship this month under Apache 2.0.
Key Takeaways
- ✓Scale: sparse MoE, 501B total / 23B active, 52 layers, interleaved local+global attention, 1M-token context
- ✓Self-reported scores: SWE-bench Verified 80.9, Terminal Bench v2.1 80.1, MCP Atlas 78.7, GPQA Diamond 90.5, AIME 2026 97.8, HLE (no tools) 36.2
- ✓RL run: 10.5K NVIDIA GB300 GPUs for four weeks, 100M+ rollouts up to 256K context, ~1.3B sandboxes, ~1M environments
- ✓Efficiency: GLM-5.2-level reasoning scores with 3–4x less inference compute; reasoning-effort parameter trades length for quality
- ✓Availability: in final red-teaming with early-access signup; Apache 2.0 weights, tech report and fine-tuning/eval stack due this month

Developer 3s Key Decision Metrics
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Reflection AI's first open-weight model, Beam (announcement), is a text-only sparse MoE with 501B total and 23B active parameters, built as an efficient workhorse for enterprise coding and agentic workloads and pitched as a Western answer to DeepSeek, Qwen and Z.ai open models.
It was pretrained on 23.8T curated tokens in under four weeks on 6,144 GB300 NVL72 GPUs (52 layers, interleaved local/global attention, fine-grained experts with near-uniform utilization), midtrained to 1M-token context, then trained with fully asynchronous high-compute RL on 10.5K GB300 GPUs for four weeks: 100M+ rollouts up to 256K context, ~1M environments, ~1.3B sandboxes, staying stable even on day-old (107-version-stale) samples. A separate safety/alignment teacher was merged in via multi-teacher on-policy distillation.
Self-reported results: SWE-bench Verified 80.9, SWE-bench Multilingual 78.0, Terminal Bench v2.1 80.1 (GLM 5.2 81.0, Kimi K3 88.3), MCP Atlas 78.7, GPQA Diamond 90.5, AIME 2026 97.8, HLE no tools 36.2. Reflection concedes Kimi K3 leads on raw capability but claims GLM-5.2-level reasoning at 3-4x less inference compute. None of this is independently verified yet.
Beam is in final red-teaming with early access via the Reflection platform. Apache 2.0 weights, tech report, model card and a run/eval/fine-tune stack are promised later in October, with distribution partners and open-source harness integrations at launch.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.