Memento 3 Unveils Model-Based Recursive Self-Improvement Using Reflective Rulebooks
agents
| Source: HF Papers | Original article
Memento 3 introduces a model‑based recursive self‑improvement framework that lets agents refine world models with reflective rulebooks when faced with limited observations.
A paper posted to arXiv on 8 October 2026 introduces **Memento 3**, a model‑based approach that lets a frozen large language model (LLM) improve its own behaviour through “reflective rulebooks”. The authors – Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou‑Ammar, Weilin Luo and Jun Wang – describe an agent that builds a natural‑language rulebook from experience, compiles that rulebook into an executable world model, and then uses the model to plan actions. Because the underlying LLM remains unchanged, the recursive self‑improvement (RSI) loop operates entirely on the mutable rulebook, sidestepping the need for costly model retraining.
The method is demonstrated on the ARC‑AGI‑3 benchmark, where a frozen‑LLM + revisable rulebook system cleared all 25 games while matching roughly 44 % of human actions. This performance marks a step forward for agents that must act in unfamiliar environments: they can infer underlying rules from sparse observations, revise those rules as new evidence arrives, and leverage the updated model to make better predictions in unseen states.
The work builds on themes we have covered recently, such as verification and self‑improvement in agentic AI (see our 9 October report) and scaling reinforcement‑learning‑based self‑improvement (MiMo‑V2.6). Memento 3 shows a concrete architecture that combines model‑based reasoning with the linguistic flexibility of LLMs, suggesting a pathway toward more robust, interpretable RSI without continual model updates.
What to watch next: replication of the results on larger, more diverse benchmarks; extensions that integrate Memento 3’s rulebooks with multi‑agent systems; and community scrutiny of safety and verification aspects as the approach moves from prototype to broader deployment.
Sources
Back to AIPULSEN