Papers
arxiv:2610.11794

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

Published on Oct 8
ยท Submitted by
Hyzhao
on Oct 9
Authors:
,
,
,
,
,
,
,

Abstract

Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.

Community

Paper submitter

Agents in unfamiliar environments must infer the underlying rules and refine their understanding through experience. We introduce Memento 3, a model-based approach to recursive self-improvement (RSI) with a frozen LLM. The agent builds a natural-language rulebook from experience, compiles it into executable code for planning, and tests and revises both against observed outcomes. The updated world model then guides further exploration and learning, creating a loop in which what the agent learns helps it decide what to investigate next.

The part I keep bumping into with self-written rules: the model can only write down the fixes it can articulate. Every reliability fix I've actually shipped this year was a schema change or a hard guard โ€” the tool now rejects a malformed arg instead of the prompt politely asking it not to. An agent that leaves itself prose notes can't do that. It's improving the layer that was never the problem, because the layer that was the problem isn't editable from inside the loop.

And when it does work, attribution is a mess. Rule fires, task passes โ€” was it the rule or the retry? I'd want the boring ablation: same run, rules disabled, does the success rate hold. If it holds, you've built an expensive diary and called it learning.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.11794
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.11794 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.11794 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.