Reinforcement learning (RL) has emerged as the defining post-training paradigm driving frontier foundation models toward autonomous self-improvement, yet scaling RL across multi-step agent environments and million-token sequences introduces profound stability and infrastructure hurdles. Xiaomi's LLM-Core Team releases the comprehensive technical report for MiMo-V2.6 (arXiv:2610.11959), an omni-modal foundation model family that scales RL compute across three critical axes. First, it employs high-throughput asynchronous training processing 1,568 samples and 2.7 to 3.7 billion tokens per optimization step across context windows scaling up to 1M tokens. Second, it orchestrates training across diverse, multi-domain environments spanning code generation, computer vision, and cybersecurity under a heterogeneous mixture of agent harnesses. Third, it implements groupwise agentic grading to mitigate reward hacking and incentivize concise, token-efficient reasoning trajectories. By freezing the MoE router and establishing multi-layered defense barriers, MiMo-V2.6 guarantees rock-solid convergence at scale, while open-sourcing its RL framework, training dynamics, and interactive environments.

Key Takeaways

  • ✓Xiaomi releases MiMo-V2.6, scaling reinforcement learning compute across 1,568 samples, 3.7B tokens per step, and up to 1M context
  • ✓Introduces groupwise agentic grading to curb reward hacking and compress trajectory token lengths by over 35%
  • ✓Freezes MoE routers to eliminate routing collapse and open-sources the complete RL post-training infrastructure
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background and the Problem

Following recent demonstrations of reinforcement learning (RL) in mathematical reasoning (e.g. DeepSeek-R1, OpenAI o-series), frontier AI labs are scaling RL post-training to omni-modal and long-horizon software engineering agents. However, extending RL into complex agentic environments introduces critical systems bottlenecks: memory exhaustion when rollouts span hundreds of thousands of tokens, reward hacking where models generate excessively verbose reasoning traces to exploit heuristic verifiers, and MoE routing collapse induced by unconstrained policy gradients.

Architecture and How It Works

Xiaomi's LLM-Core Team introduces MiMo-V2.6 (arXiv:2610.11959), detailing a scalable infrastructure for agentic reinforcement learning:

  1. 3D Scaling of RL Compute:
  • Asynchronous High-Throughput Rollouts: Consumes 1,568 samples and 2.7 to 3.7 billion tokens per optimization step, pushing stable post-training context windows to 1M tokens.
  • Multi-Domain Agent Environments: Integrates diverse sandbox environments spanning code development, cyber defense, and visual manipulation managed by heterogeneous agent harnesses.
  • Groupwise Agentic Grading: Deploys peer agent grading panels to evaluate long-horizon trajectories, penalizing bloated reasoning to steer models toward concise, token-efficient code.
  1. Router Freezing & Anti-Hacking Guardrails: Freezes MoE routing gates during RL post-training to prevent expert capacity collapse, accompanied by multi-stage reward sanitization pipelines.
  2. Open Ecosystem Contribution: Open-sources the complete RL framework, training trajectories, and interactive environments to democratize frontier self-improvement research.

Benchmarks and Measured Results

Benchmarked across demanding coding, reasoning, and multi-modal benchmarks:

  1. Substantial Resolve Rate Gains: Delivers sharp accuracy improvements over baseline SFT checkpoints across SWE-bench variants and deep agent environments.
  2. 35% Token Overhead Reduction: Groupwise grading compresses average solution lengths by over 35% while matching or exceeding the solution quality of unconstrained reasoning paths.
  3. Million-Token Training Stability: Maintains monotonic loss and reward convergence across sequence lengths scaling up to 1M tokens without gradient divergence.

Getting Started for Developers

The MiMo-V2.6 technical report establishes scalable engineering patterns for post-training reasoning models. Platform architects deploying RL infrastructure should decouple environment simulation from gradient optimization, lock down MoE routing weights to ensure representation stability, and integrate token-efficiency penalties into automated reward functions to minimize inference costs during enterprise deployment.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.