ahmedehabb commited on
Commit
80df427
·
verified ·
1 Parent(s): eb75b77

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +57 -0
README.md ADDED
@@ -0,0 +1,57 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-7B-Instruct
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - memory
7
+ - long-horizon
8
+ - reinforcement-learning
9
+ - grpo
10
+ - agent
11
+ ---
12
+
13
+ # Memory-R2 7B — Memory Manager (LoGo-GRPO, 32-session champion)
14
+
15
+ This is the trained **memory-management policy** from [Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents](https://arxiv.org/abs/2605.21768) (arXiv:2605.21768).
16
+
17
+ It is a Qwen2.5-7B-Instruct model fine-tuned with **LoGo-GRPO** (turn-level + token-level credit assignment) via a curriculum of 8 → 16 → 32-session rollouts on the LoCoMo long-horizon dialogue dataset. Given a running conversation, it decides what to INSERT / UPDATE / DELETE in an external memory store, so that a separate answer agent can later answer questions using only the maintained memory.
18
+
19
+ This checkpoint is the paper's deployed "champion" (`32sess_champion_v2`, LoGo-GRPO curriculum, global step 5) — the memory manager behind the headline `tab:main` **Memory-R2 (OSS)** result on LoCoMo:
20
+
21
+ | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
22
+ | --- | ---: | ---: | ---: |
23
+ | GPT-OSS-120B (untrained, paired at eval time) | **49.29** | **43.64** | **86.08** |
24
+
25
+ See the paper's `tab:different-answer-agent` for how this same memory manager performs when paired with other answer agents (untrained Qwen-7B, RL-trained Qwen-7B, etc.).
26
+
27
+ ## Usage
28
+
29
+ This model is one half of a two-agent pipeline (memory manager + answer agent) and is not intended as a general-purpose chat model. Full inference code, the memory-store protocol, and the answer-agent pairing are in the [project repository](https://github.com/) (see the paper for the official release).
30
+
31
+ ```python
32
+ from transformers import AutoModelForCausalLM, AutoTokenizer
33
+
34
+ model = AutoModelForCausalLM.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="memory-manager", torch_dtype="auto", device_map="auto")
35
+ tokenizer = AutoTokenizer.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="memory-manager")
36
+ ```
37
+
38
+ ## Training
39
+
40
+ - Base model: `Qwen/Qwen2.5-7B-Instruct`
41
+ - Algorithm: LoGo-GRPO (turn-level + token-level advantage), curriculum-trained 8-session → 16-session → 32-session
42
+ - Reward: per-session cumulative F1 against gold QA + a memory-compression penalty (λ=0.3)
43
+ - Judge for reward/logging during training: GPT-OSS-120B
44
+
45
+ ## Citation
46
+
47
+ ```bibtex
48
+ @misc{yan2026memoryr2faircreditassignment,
49
+ title={Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents},
50
+ author={Sikuan Yan and Ahmed Bahloul and Ercong Nie and Susanna Schwarzmann and Riccardo Trivisonno and Volker Tresp and Yunpu Ma},
51
+ year={2026},
52
+ eprint={2605.21768},
53
+ archivePrefix={arXiv},
54
+ primaryClass={cs.LG},
55
+ url={https://arxiv.org/abs/2605.21768},
56
+ }
57
+ ```