ahmedehabb commited on
Commit
ca55194
Β·
verified Β·
1 Parent(s): 718e5d0

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +17 -22
README.md CHANGED
@@ -10,46 +10,41 @@ tags:
10
  - agent
11
  ---
12
 
13
- # Memory-R2 7B β€” Memory Manager + Answer Agent
14
 
15
- Checkpoints from [Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents](https://arxiv.org/abs/2605.21768) (arXiv:2605.21768).
16
 
17
- This repo contains **two separate Qwen2.5-7B-Instruct checkpoints** that play two different, non-overlapping roles in the pipeline. They are not interchangeable and neither one does the other's job:
18
 
19
- | Subfolder | Role | Does it manage memory? | Does it answer questions? |
20
- | --- | --- | :---: | :---: |
21
- | **`memory-manager/`** | Stage 1 β€” reads the running conversation and decides what to INSERT / UPDATE / DELETE in the external memory store. | βœ… Yes β€” this is its only job | ❌ No β€” it never sees or answers the held-out QA questions |
22
- | **`answer-agent/`** | Stage 2 β€” given a question and the memory store `memory-manager/` produced, generates the final answer. | ❌ No β€” it never touches the memory store's write operations | βœ… Yes β€” this is its only job |
23
-
24
- **`memory-manager/`** is `32sess_champion_v2` β€” the LoGo-GRPO-trained (turn-level + token-level credit assignment, curriculum 8β†’16β†’32 sessions on LoCoMo) memory-management policy, and is the paper's main contribution / "the champion" checkpoint. It is the piece you need regardless of which answer agent you pair it with.
25
-
26
- **`answer-agent/`** (`sft_cont_step55`, SFT+RL-trained) is an *optional, swappable* stage-2 model β€” the memory manager was evaluated in the paper against several different answer agents, and this is simply the best-performing one we trained ourselves. You can equally pair `memory-manager/` with an untrained Qwen-7B, GPT-OSS-120B, or any other instruction-tuned LLM as the answer agent β€” see `tab:different-answer-agent` in the paper. The reverse never happens: `answer-agent/` is not used for memory operations, and swapping it out has no effect on how memory is maintained.
27
 
28
  ## Headline results (`tab:main`)
29
 
30
- `memory-manager/` is the constant in every row below; only the answer agent changes:
31
 
32
- | Memory manager | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
33
- | --- | --- | ---: | ---: | ---: |
34
- | memory-manager/ | `answer-agent/` (ours, SFT+RL) | **51.46** | **44.84** | 69.03 |
35
- | memory-manager/ | GPT-OSS-120B (untrained, external) | 49.29 | 43.64 | **86.08** |
36
 
37
- ## Usage
38
 
39
- These models are two halves of a two-agent pipeline (memory manager + answer agent) and are not intended as general-purpose chat models on their own. Full inference code and the memory-store protocol are in the [project repository](https://github.com/) (see the paper for the official release).
40
 
41
  ```python
42
  from transformers import AutoModelForCausalLM, AutoTokenizer
43
 
44
- memory_manager = AutoModelForCausalLM.from_pretrained("ahmedehabb/Memory-R2", subfolder="memory-manager", torch_dtype="auto", device_map="auto")
45
- answer_agent = AutoModelForCausalLM.from_pretrained("ahmedehabb/Memory-R2", subfolder="answer-agent", torch_dtype="auto", device_map="auto")
46
  ```
47
 
 
 
48
  ## Training
49
 
50
  - Base model: `Qwen/Qwen2.5-7B-Instruct`
51
- - Memory manager: LoGo-GRPO (turn-level + token-level advantage), curriculum-trained 8-session β†’ 16-session β†’ 32-session; reward = per-session cumulative F1 against gold QA + a memory-compression penalty (Ξ»=0.3)
52
- - Answer agent: SFT warm-start followed by an RL continuation (answer-F1 reward) against the memory manager's rollouts
53
  - Judge for reward/logging during training: GPT-OSS-120B
54
 
55
  ## Citation
 
10
  - agent
11
  ---
12
 
13
+ # Memory-R2 7B β€” Memory Manager
14
 
15
+ The trained **memory-management policy** from [Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents](https://arxiv.org/abs/2605.21768) (arXiv:2605.21768). This is the paper's main contribution and deployed "champion" (`32sess_champion_v2`, LoGo-GRPO curriculum, global step 5).
16
 
17
+ It is a Qwen2.5-7B-Instruct model fine-tuned with **LoGo-GRPO** (turn-level + token-level credit assignment) via a curriculum of 8 β†’ 16 β†’ 32-session rollouts on the LoCoMo long-horizon dialogue dataset. Given a running conversation, it decides what to INSERT / UPDATE / DELETE in an external memory store.
18
 
19
+ **This model only manages memory β€” it does not answer questions.** A separate answer agent reads the memory store this model produces and generates answers; it can be any instruction-tuned LLM. Our own SFT+RL-trained answer agent is released separately at **[ahmedehabb/Memory-R2-answer-agent](https://huggingface.co/ahmedehabb/Memory-R2-answer-agent)**.
 
 
 
 
 
 
 
20
 
21
  ## Headline results (`tab:main`)
22
 
23
+ This memory manager is held constant; only the paired answer agent changes:
24
 
25
+ | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
26
+ | --- | ---: | ---: | ---: |
27
+ | [ahmedehabb/Memory-R2-answer-agent](https://huggingface.co/ahmedehabb/Memory-R2-answer-agent) (ours, SFT+RL) | **51.46** | **44.84** | 69.03 |
28
+ | GPT-OSS-120B (untrained, external) | 49.29 | 43.64 | **86.08** |
29
 
30
+ See the paper's `tab:different-answer-agent` for more pairings (untrained Qwen-7B, etc.) β€” the memory manager is not tied to any one answer agent.
31
 
32
+ ## Usage
33
 
34
  ```python
35
  from transformers import AutoModelForCausalLM, AutoTokenizer
36
 
37
+ memory_manager = AutoModelForCausalLM.from_pretrained("ahmedehabb/Memory-R2", torch_dtype="auto", device_map="auto")
38
+ tokenizer = AutoTokenizer.from_pretrained("ahmedehabb/Memory-R2")
39
  ```
40
 
41
+ Full inference code and the memory-store protocol are in the [project repository](https://github.com/ahmedehabb/Memory-R2) (see the paper for the official release).
42
+
43
  ## Training
44
 
45
  - Base model: `Qwen/Qwen2.5-7B-Instruct`
46
+ - Algorithm: LoGo-GRPO (turn-level + token-level advantage), curriculum-trained 8-session β†’ 16-session β†’ 32-session
47
+ - Reward: per-session cumulative F1 against gold QA + a memory-compression penalty (Ξ»=0.3)
48
  - Judge for reward/logging during training: GPT-OSS-120B
49
 
50
  ## Citation