Vals released Vibe Code Bench 1-100 (VCB 1-100), which starts from a fully passing web app and asks models to apply a long sequence of dependent product requests without breaking prior behavior; the first incomplete iteration ends the run. On the published leaderboard Claude Opus 5 leads at 28.53%, with Claude Fable 5.1 at 28.00% and GPT-6 Astra at 27.64%. Unlike one-shot greenfield coding benches, VCB 1-100 targets sustained multi-turn evolution typical of vibe coding and agentic software work.
Key Takeaways
- βVCB 1-100 starts from a fully passing app and ends on the first incomplete dependent change.
- βLeaderboard: Claude Opus 5 28.53%, Claude Fable 5.1 28.00%, GPT-6 Astra 27.64%.
- βVals notes self-testing gaps: ~1/3 of Opus 5 and Fable 5.1 tool calls are browser use; Grok 4.6 almost never opens a browser.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.