Anthropic released Fellows research showing Claude can research, propose methods, then train and test smaller models to fix alignment failures—given 48 hours and one GPU. Across 10 failure types it improved safety scores without hurting general capabilities, and the best methods generalized to held-out benchmarks, the Petri behavioral audit, and models up to 4.7x larger. In a first successor-alignment test, Sonnet 5 post-trained an early Opus 4.8 checkpoint to safety scores approaching production Opus 4.8.

Key Takeaways

  • Autonomous alignment is now a reproducible train-eval loop: Claude hill-climbs safety while preserving capabilities.
  • Sonnet 5 post-training an Opus 4.8 checkpoint is early evidence a weaker model can align a stronger successor.
  • Anthropic is releasing the automated alignment research setup; rare failures still depend on measuring the right things.
ADSponsored