Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -24,22 +24,31 @@ This checkpoint is the paper's deployed "champion" (`32sess_champion_v2`, LoGo-G
|
|
| 24 |
|
| 25 |
See the paper's `tab:different-answer-agent` for how this same memory manager performs when paired with other answer agents (untrained Qwen-7B, RL-trained Qwen-7B, etc.).
|
| 26 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
## Usage
|
| 28 |
|
| 29 |
-
|
| 30 |
|
| 31 |
```python
|
| 32 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 33 |
|
| 34 |
-
|
| 35 |
-
|
| 36 |
```
|
| 37 |
|
| 38 |
## Training
|
| 39 |
|
| 40 |
- Base model: `Qwen/Qwen2.5-7B-Instruct`
|
| 41 |
-
-
|
| 42 |
-
-
|
| 43 |
- Judge for reward/logging during training: GPT-OSS-120B
|
| 44 |
|
| 45 |
## Citation
|
|
|
|
| 24 |
|
| 25 |
See the paper's `tab:different-answer-agent` for how this same memory manager performs when paired with other answer agents (untrained Qwen-7B, RL-trained Qwen-7B, etc.).
|
| 26 |
|
| 27 |
+
## Answer agent (`answer-agent/`)
|
| 28 |
+
|
| 29 |
+
Also included is `sft_cont_step55`, a separately SFT+RL-trained Qwen2.5-7B-Instruct **answer agent**: given the memory manager's maintained memory store, it answers the held-out questions. Paired with the `memory-manager/` checkpoint above, this is the paper's highest-F1 configuration ("Memory-R2 (SFT-RL)" in `tab:main`):
|
| 30 |
+
|
| 31 |
+
| Memory manager | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
|
| 32 |
+
| --- | --- | ---: | ---: | ---: |
|
| 33 |
+
| memory-manager/ | answer-agent/ | **51.46** | **44.84** | 69.03 |
|
| 34 |
+
| memory-manager/ | GPT-OSS-120B (untrained) | 49.29 | 43.64 | **86.08** |
|
| 35 |
+
|
| 36 |
## Usage
|
| 37 |
|
| 38 |
+
These models are two halves of a two-agent pipeline (memory manager + answer agent) and are not intended as general-purpose chat models on their own. Full inference code and the memory-store protocol are in the [project repository](https://github.com/) (see the paper for the official release).
|
| 39 |
|
| 40 |
```python
|
| 41 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 42 |
|
| 43 |
+
memory_manager = AutoModelForCausalLM.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="memory-manager", torch_dtype="auto", device_map="auto")
|
| 44 |
+
answer_agent = AutoModelForCausalLM.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="answer-agent", torch_dtype="auto", device_map="auto")
|
| 45 |
```
|
| 46 |
|
| 47 |
## Training
|
| 48 |
|
| 49 |
- Base model: `Qwen/Qwen2.5-7B-Instruct`
|
| 50 |
+
- Memory manager: LoGo-GRPO (turn-level + token-level advantage), curriculum-trained 8-session → 16-session → 32-session; reward = per-session cumulative F1 against gold QA + a memory-compression penalty (λ=0.3)
|
| 51 |
+
- Answer agent: SFT warm-start followed by an RL continuation (answer-F1 reward) against the memory manager's rollouts
|
| 52 |
- Judge for reward/logging during training: GPT-OSS-120B
|
| 53 |
|
| 54 |
## Citation
|