Developer 3s Key Decision Metrics
On 2026-10-01 Ai2 released Olmo-core 3, an open MoE training stack (GitHub, PyPI ai2-olmo-core). It moves from FSDP to DDP + expert parallelism; on 8×B300 a 47B MoE hits ~2.7× throughput (52k vs 19.4k tok/s/GPU). Systems benches include 1.2T on 512 GPUs at up to 858 useful-model TFLOP/s/GPU and a DeepEP v2 capacity probe at 2.38T. Report: olmocore3.
Key Takeaways
- ✓Announced 2026-10-01: allenai.org/blog/olmocore3 + HF mirror + github.com/allenai/OLMo-core
- ✓Stack: EP/PP/distributed optimizer; NVSHMEM rowwise EP + GPU-resident routing + grouped GEMM; optional DeepEP v2
- ✓Throughput: expert pool 8→128 at ~3.2B active, <5% drop; 47B MoE on 8×B300 52k vs 19.4k tok/s/GPU (~2.7×)
- ✓Scale: 1.2T / 58.36B active / 512 GPUs up to 858 TFLOP/s/GPU (random-routing systems ref); DeepEP v2 probe 2.38T
- ✓MXFP8 ~+21% vs BF16, peak active mem 103→95 GiB; topology-agnostic checkpoints; next Olmo will be MoE
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Core Background & Industry Pain Points
Frontier MoEs grow total capacity far faster than per-token active compute, but dense-era full-reshard FSDP still gathers weights by total expert capacity each microbatch—Ai2’s “capacity tax.” Labs need an open stack that still works near trillion scale. On 2026-10-01 Ai2 shipped Olmo-core 3 (HF mirror) as infrastructure for the next MoE Olmo.
Architecture Highlights & Internals
The new path centers on DDP: keep experts resident and route token rows. Compose EP, PP, and a distributed optimizer. Systems pieces: NVSHMEM rowwise EP, GPU-resident routing, device-scheduled grouped GEMM, MXFP8, topology-agnostic FP32 checkpoints, optional DeepEP v2. Code: allenai/OLMo-core / PyPI ai2-olmo-core. Report: olmocore3.
Authoritative Benchmarks & Measured Scores
Figures are Ai2 systems numbers under random routing—not model quality. Capacity sweep: top-4, ~3.2B active, experts 8→128, throughput drop <5%. 8×B300 47B MoE: ~52k vs ~19.4k tok/s/GPU (~2.7×). 1.2T / 58.36B active / 512 GPUs up to 858 TFLOP/s/GPU; DeepEP v2 probe 2.38T. MXFP8: ~+21% vs BF16, peak active mem 103→95 GiB. Report also documents token gerrymandering and failed overlap tricks—don’t treat as a Megatron replacement claim.
Developer Hands-on Guide
Read blog + report §14 conditions; install ai2-olmo-core or clone GitHub; validate EP/PP/MXFP8/recompute on small runs before scaling; compare with Megatron-Core on your topology; this is pretrain infra, not an inference agent model.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.