@huangruiteng reported exploratory LHTB results for LoopX 1.0.3 Heartbeat with GPT-5.6-sol: mean reward 0.4948, above Claude Fable 5’s published 0.487 reference and about 10.6% above native Codex Goal (0.4475). LHTB spans nine long-horizon categories beyond coding; LoopX keeps work state outside the session and wakes executors via Heartbeat with plan checks. The author notes single-run-per-task arms and different cost accounting, so same-budget efficiency is not claimed.
Key Takeaways
- ✓GPT-5.6-sol + LoopX 1.0.3 hits 0.4948 mean on LHTB, above Claude Fable 5’s 0.487 reference and Codex Goal’s 0.4475.
- ✓Vs Goal: 23 wins / 13 ties / 10 losses; tasks with reward ≥0.95 rise from 4/46 to 7/46.
- ✓LoopX uses out-of-session state + Heartbeat executors; results are exploratory with differing cost accounting.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.