Large language model post-training generates vast self-generated rollouts through reinforcement learning and on-policy distillation, yet these trajectories are conventionally discarded as stale once policy weights advance. However, historical rollouts preserve valuable exploratory behaviors that newer policies no longer express reliably, despite containing abandoned branches and syntax dead-ends. Researchers introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves full historical trajectories as context while applying loss strictly to verified model-generated continuations. Without running a single additional rollout, ROSS elevates Qwen3.6-35B-A3B's MOPD six-benchmark average from 58.40% to 62.20% and surges SWE-bench Verified from 64.20% to 68.40%, unlocking a highly cost-effective paradigm for agent post-training.
Key Takeaways
- ✓Recycling 'Stale' Rollout Experience: Challenges the conventional assumption that historical agent rollouts become obsolete as policy weights evolve, establishing an offline recycling mechanism.
- ✓Selective Supervision Filtering: Preserves the entire multi-step interaction history as prompt context while restricting loss backpropagation exclusively to verified, high-quality continuation segments.
- ✓68.40% on SWE-bench Verified (+4.20% Gain): Applied to Qwen3.6-35B-A3B without generating new rollouts, ROSS pushes SWE-bench Verified from 64.20% to 68.40% and elevates MOPD average benchmarks by 3.8 percentage points.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.