Clarify two-checkpoint roles: memory-manager vs answer-agent
Browse files
README.md
CHANGED
|
@@ -10,28 +10,29 @@ tags:
|
|
| 10 |
- agent
|
| 11 |
---
|
| 12 |
|
| 13 |
-
# Memory-R2 7B β Memory Manager
|
| 14 |
|
| 15 |
-
|
| 16 |
|
| 17 |
-
|
| 18 |
|
| 19 |
-
|
|
|
|
|
|
|
|
|
|
| 20 |
|
| 21 |
-
|
| 22 |
-
| --- | ---: | ---: | ---: |
|
| 23 |
-
| GPT-OSS-120B (untrained, paired at eval time) | **49.29** | **43.64** | **86.08** |
|
| 24 |
|
| 25 |
-
|
| 26 |
|
| 27 |
-
##
|
| 28 |
|
| 29 |
-
|
| 30 |
|
| 31 |
| Memory manager | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
|
| 32 |
| --- | --- | ---: | ---: | ---: |
|
| 33 |
-
| memory-manager/ | answer-agent/ | **51.46** | **44.84** | 69.03 |
|
| 34 |
-
| memory-manager/ | GPT-OSS-120B (untrained) | 49.29 | 43.64 | **86.08** |
|
| 35 |
|
| 36 |
## Usage
|
| 37 |
|
|
|
|
| 10 |
- agent
|
| 11 |
---
|
| 12 |
|
| 13 |
+
# Memory-R2 7B β Memory Manager + Answer Agent
|
| 14 |
|
| 15 |
+
Checkpoints from [Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents](https://arxiv.org/abs/2605.21768) (arXiv:2605.21768).
|
| 16 |
|
| 17 |
+
This repo contains **two separate Qwen2.5-7B-Instruct checkpoints** that play two different, non-overlapping roles in the pipeline. They are not interchangeable and neither one does the other's job:
|
| 18 |
|
| 19 |
+
| Subfolder | Role | Does it manage memory? | Does it answer questions? |
|
| 20 |
+
| --- | --- | :---: | :---: |
|
| 21 |
+
| **`memory-manager/`** | Stage 1 β reads the running conversation and decides what to INSERT / UPDATE / DELETE in the external memory store. | β
Yes β this is its only job | β No β it never sees or answers the held-out QA questions |
|
| 22 |
+
| **`answer-agent/`** | Stage 2 β given a question and the memory store `memory-manager/` produced, generates the final answer. | β No β it never touches the memory store's write operations | β
Yes β this is its only job |
|
| 23 |
|
| 24 |
+
**`memory-manager/`** is `32sess_champion_v2` β the LoGo-GRPO-trained (turn-level + token-level credit assignment, curriculum 8β16β32 sessions on LoCoMo) memory-management policy, and is the paper's main contribution / "the champion" checkpoint. It is the piece you need regardless of which answer agent you pair it with.
|
|
|
|
|
|
|
| 25 |
|
| 26 |
+
**`answer-agent/`** (`sft_cont_step55`, SFT+RL-trained) is an *optional, swappable* stage-2 model β the memory manager was evaluated in the paper against several different answer agents, and this is simply the best-performing one we trained ourselves. You can equally pair `memory-manager/` with an untrained Qwen-7B, GPT-OSS-120B, or any other instruction-tuned LLM as the answer agent β see `tab:different-answer-agent` in the paper. The reverse never happens: `answer-agent/` is not used for memory operations, and swapping it out has no effect on how memory is maintained.
|
| 27 |
|
| 28 |
+
## Headline results (`tab:main`)
|
| 29 |
|
| 30 |
+
`memory-manager/` is the constant in every row below; only the answer agent changes:
|
| 31 |
|
| 32 |
| Memory manager | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
|
| 33 |
| --- | --- | ---: | ---: | ---: |
|
| 34 |
+
| memory-manager/ | `answer-agent/` (ours, SFT+RL) | **51.46** | **44.84** | 69.03 |
|
| 35 |
+
| memory-manager/ | GPT-OSS-120B (untrained, external) | 49.29 | 43.64 | **86.08** |
|
| 36 |
|
| 37 |
## Usage
|
| 38 |
|