Endowing artificial agents with general intelligence requires inferring environment dynamics and refining world models continuously in the wild. Fine-tuning models directly is prohibitively expensive and precipitates catastrophic forgetting. Researchers from University College London (UCL) and Huawei Noah's Ark Lab UK present Memento 3 (arXiv:2610.11794), a framework enabling frozen LLM agents to achieve model-based Recursive Self-Improvement (RSI) purely via external memory. The agent maintains a natural-language rulebook representing its evolving hypotheses of environment physics, which is compiled into executable code for fast simulation and action planning. Governed by a continuous observation-reflection-revision-compilation-verification loop, the system leverages prediction errors to refine rules. Edits are merged only after passing LLM semantic alignment audits and deterministic cell-exact replay verification. On the rigorous ARC-AGI-3 abstraction benchmark, Memento 3 solves 100% of levels across all 25 public environments, scoring a perfect 100.0 Relative Human Action Efficiency (RHAE) while utilizing only 44% of human action counts. In Atari Pong, a synthesized feedback controller shuts out opponents 21:0 across three evaluation seeds with zero test-time LLM inference.
Key Takeaways
- ✓UCL and Huawei Noah's Ark introduce Memento 3, unlocking model-based Recursive Self-Improvement for frozen LLMs via executable rulebooks
- ✓Compiles natural-language physics hypotheses into code world models validated through cell-exact replay verification
- ✓Clears all 25 games on ARC-AGI-3 using only 44% of human actions, and delivers 21:0 Pong shutouts with zero runtime LLM calls
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Background and the Problem
Achieving Recursive Self-Improvement (RSI) remains a core grand challenge in artificial intelligence. When deployed in novel task environments, agents must deduce underlying dynamics and adjust their policies based on empirical observations. Modifying large language model (LLM) weights via continual fine-tuning triggers catastrophic forgetting and carries prohibitive training costs. Conversely, purely textual episodic memories lack formal predictive rigor, failing to support precise mental simulation or multi-step forward planning in challenging abstraction tasks.
Architecture and How It Works
Researchers from University College London (UCL) and Huawei Noah's Ark Lab UK introduce Memento 3 (arXiv:2610.11794), realizing model-based RSI on frozen models:
- Rulebook-to-Code Compilation: Maintains external semantic memory as an interpretable natural-language rulebook documenting environmental physics and mechanics, which is compiled into executable simulation code for low-latency planning.
- Cell-Exact Replay Verification: Employs an observation-reflection-revision-compilation-verification loop. Candidate model modifications are accepted only after passing deterministic cell-exact replay of historical state transitions alongside LLM semantic alignment audits.
- Population-Based Exploratory Co-design: Scales to multi-agent populations maintaining diverse world model candidates, utilizing epistemic disagreement to guide active exploration.
- Parameter-Free RSI: Delivers recursive capability accumulation entirely through program synthesis and verification while keeping foundational LLM parameters static.
Benchmarks and Measured Results
Benchmarked on demanding reasoning and continuous control benchmarks:
- 100% Clearance on ARC-AGI-3: Memento 3 clears all levels across all 25 public games in the ARC-AGI-3 benchmark, registering a perfect Relative Human Action Efficiency (RHAE) score of 100.0.
- 44% Human Action Budget: Solves complex procedural abstraction problems using only 44% of the physical action counts consumed by human benchmark subjects.
- 21:0 Pong Shutout with Zero Inference Calls: In Atari Pong, synthesizes an executable closed-loop controller that achieves 21:0 shutouts across three varied opening seeds without triggering a single additional LLM token query.
Getting Started for Developers
Memento 3 offers an architectural blueprint for embodied autonomy, automated scientific discovery, and proprietary system reverse-engineering. Rather than treating foundation models as end-to-end controllers, systems engineers should configure LLMs as meta-programmers. Generating and verifying executable code world models provides deterministic simulation, zero-latency inference, and verifiable self-improvement without parameter fine-tuning.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.