Large language model post-training generates vast self-generated rollouts through reinforcement learning and on-policy distillation, yet these trajectories are conventionally discarded as stale once policy weights advance. However, historical rollouts preserve valuable exploratory behaviors that newer policies no longer express reliably, despite containing abandoned branches and syntax dead-ends. Researchers introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves full historical trajectories as context while applying loss strictly to verified model-generated continuations. Without running a single additional rollout, ROSS elevates Qwen3.6-35B-A3B's MOPD six-benchmark average from 58.40% to 62.20% and surges SWE-bench Verified from 64.20% to 68.40%, unlocking a highly cost-effective paradigm for agent post-training.

Key Takeaways

  • ✓Recycling 'Stale' Rollout Experience: Challenges the conventional assumption that historical agent rollouts become obsolete as policy weights evolve, establishing an offline recycling mechanism.
  • ✓Selective Supervision Filtering: Preserves the entire multi-step interaction history as prompt context while restricting loss backpropagation exclusively to verified, high-quality continuation segments.
  • ✓68.40% on SWE-bench Verified (+4.20% Gain): Applied to Qwen3.6-35B-A3B without generating new rollouts, ROSS pushes SWE-bench Verified from 64.20% to 68.40% and elevates MOPD average benchmarks by 3.8 percentage points.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points Post-training workflows for frontier coding agents and reasoning LLMs rely heavily on online reinforcement learning (RL) and on-policy distillation. However, engineering teams face massive compute wastage in production clusters: 1. Disposability of High-Cost Rollouts: Generating millions of interactive rollout trajectories across containerized environments consumes thousands of GPU hours. Under standard RL protocols, once the policy distribution shifts, historical trajectories are deemed stale and discarded; 2. Noise Pollution in Naive SFT Recycling: Directly feeding raw historical rollouts into supervised fine-tuning (SFT) backfires: the rollouts are riddled with exploratory missteps, infinite loops, and redundant shell commands that poison policy behavior if blindly imitated. ### Architectural Highlights & Underlying Mechanics The authors introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), unlocking hidden value in archived trajectory logs: 1. Full Context Preservation with Targeted Continuation Loss: Preserves the complete historical interactive trace as prompt context—ensuring the model perceives realistic error recovery states—while applying backpropagation loss strictly to selected high-yield continuation tokens; 2. Finer-Grained Selective Masking: Evaluates intermediate execution transitions against outcome verifiers and AST diffs, filtering out failed trials and dead-end attempts; 3. Sampling-Free Secondary Refinement: Bypasses the latency and GPU overhead of spinning up fresh interactive environments, operating directly on persisted trajectory logs via lightweight offline SFT. ### Benchmark & Experimental Validation - SWE-bench Verified Surges to 68.40%: On the benchmark for real-world software engineering issues, applying ROSS to Qwen3.6-35B-A3B pushes resolution rates from 64.20% to 68.40% (a +4.20% absolute gain); - Broad Gains Across MOPD Benchmarks: Evaluated on multi-teacher on-policy distillation across mathematics, code synthesis, and instruction following, the six-benchmark average increases from 58.40% to 62.20%; - Zero Additional Rollout Overhead: Eliminates the expensive rollout generation phase, achieving substantial downstream gains at less than 7% of the compute budget required for an equivalent online RL pass. ### Engineering Takeaways & Practical Guide - Paper & Reference: Full formal formulations and masking strategies are documented in arXiv:2609.35954; - Operational Recipe for AI Labs: Teams running agent RL post-training should persist all rollout buffers to object storage. Running a final selective-supervision pass over these trajectories post-RL reliably consolidates reasoning quality without spending compute on fresh rollouts; - Cross-Domain Utility: Applicable beyond code to multi-step math problem solving and interactive tool use.