ahmedehabb commited on
Commit
489e989
Β·
verified Β·
1 Parent(s): c76f0c0

Clarify two-checkpoint roles: memory-manager vs answer-agent

Browse files
Files changed (1) hide show
  1. README.md +13 -12
README.md CHANGED
@@ -10,28 +10,29 @@ tags:
10
  - agent
11
  ---
12
 
13
- # Memory-R2 7B β€” Memory Manager (LoGo-GRPO, 32-session champion)
14
 
15
- This is the trained **memory-management policy** from [Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents](https://arxiv.org/abs/2605.21768) (arXiv:2605.21768).
16
 
17
- It is a Qwen2.5-7B-Instruct model fine-tuned with **LoGo-GRPO** (turn-level + token-level credit assignment) via a curriculum of 8 β†’ 16 β†’ 32-session rollouts on the LoCoMo long-horizon dialogue dataset. Given a running conversation, it decides what to INSERT / UPDATE / DELETE in an external memory store, so that a separate answer agent can later answer questions using only the maintained memory.
18
 
19
- This checkpoint is the paper's deployed "champion" (`32sess_champion_v2`, LoGo-GRPO curriculum, global step 5) β€” the memory manager behind the headline `tab:main` **Memory-R2 (OSS)** result on LoCoMo:
 
 
 
20
 
21
- | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
22
- | --- | ---: | ---: | ---: |
23
- | GPT-OSS-120B (untrained, paired at eval time) | **49.29** | **43.64** | **86.08** |
24
 
25
- See the paper's `tab:different-answer-agent` for how this same memory manager performs when paired with other answer agents (untrained Qwen-7B, RL-trained Qwen-7B, etc.).
26
 
27
- ## Answer agent (`answer-agent/`)
28
 
29
- Also included is `sft_cont_step55`, a separately SFT+RL-trained Qwen2.5-7B-Instruct **answer agent**: given the memory manager's maintained memory store, it answers the held-out questions. Paired with the `memory-manager/` checkpoint above, this is the paper's highest-F1 configuration ("Memory-R2 (SFT-RL)" in `tab:main`):
30
 
31
  | Memory manager | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
32
  | --- | --- | ---: | ---: | ---: |
33
- | memory-manager/ | answer-agent/ | **51.46** | **44.84** | 69.03 |
34
- | memory-manager/ | GPT-OSS-120B (untrained) | 49.29 | 43.64 | **86.08** |
35
 
36
  ## Usage
37
 
 
10
  - agent
11
  ---
12
 
13
+ # Memory-R2 7B β€” Memory Manager + Answer Agent
14
 
15
+ Checkpoints from [Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents](https://arxiv.org/abs/2605.21768) (arXiv:2605.21768).
16
 
17
+ This repo contains **two separate Qwen2.5-7B-Instruct checkpoints** that play two different, non-overlapping roles in the pipeline. They are not interchangeable and neither one does the other's job:
18
 
19
+ | Subfolder | Role | Does it manage memory? | Does it answer questions? |
20
+ | --- | --- | :---: | :---: |
21
+ | **`memory-manager/`** | Stage 1 β€” reads the running conversation and decides what to INSERT / UPDATE / DELETE in the external memory store. | βœ… Yes β€” this is its only job | ❌ No β€” it never sees or answers the held-out QA questions |
22
+ | **`answer-agent/`** | Stage 2 β€” given a question and the memory store `memory-manager/` produced, generates the final answer. | ❌ No β€” it never touches the memory store's write operations | βœ… Yes β€” this is its only job |
23
 
24
+ **`memory-manager/`** is `32sess_champion_v2` β€” the LoGo-GRPO-trained (turn-level + token-level credit assignment, curriculum 8β†’16β†’32 sessions on LoCoMo) memory-management policy, and is the paper's main contribution / "the champion" checkpoint. It is the piece you need regardless of which answer agent you pair it with.
 
 
25
 
26
+ **`answer-agent/`** (`sft_cont_step55`, SFT+RL-trained) is an *optional, swappable* stage-2 model β€” the memory manager was evaluated in the paper against several different answer agents, and this is simply the best-performing one we trained ourselves. You can equally pair `memory-manager/` with an untrained Qwen-7B, GPT-OSS-120B, or any other instruction-tuned LLM as the answer agent β€” see `tab:different-answer-agent` in the paper. The reverse never happens: `answer-agent/` is not used for memory operations, and swapping it out has no effect on how memory is maintained.
27
 
28
+ ## Headline results (`tab:main`)
29
 
30
+ `memory-manager/` is the constant in every row below; only the answer agent changes:
31
 
32
  | Memory manager | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
33
  | --- | --- | ---: | ---: | ---: |
34
+ | memory-manager/ | `answer-agent/` (ours, SFT+RL) | **51.46** | **44.84** | 69.03 |
35
+ | memory-manager/ | GPT-OSS-120B (untrained, external) | 49.29 | 43.64 | **86.08** |
36
 
37
  ## Usage
38