ahmedehabb commited on
Commit
02d8ae5
·
verified ·
1 Parent(s): 0bb0624

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +14 -5
README.md CHANGED
@@ -24,22 +24,31 @@ This checkpoint is the paper's deployed "champion" (`32sess_champion_v2`, LoGo-G
24
 
25
  See the paper's `tab:different-answer-agent` for how this same memory manager performs when paired with other answer agents (untrained Qwen-7B, RL-trained Qwen-7B, etc.).
26
 
 
 
 
 
 
 
 
 
 
27
  ## Usage
28
 
29
- This model is one half of a two-agent pipeline (memory manager + answer agent) and is not intended as a general-purpose chat model. Full inference code, the memory-store protocol, and the answer-agent pairing are in the [project repository](https://github.com/) (see the paper for the official release).
30
 
31
  ```python
32
  from transformers import AutoModelForCausalLM, AutoTokenizer
33
 
34
- model = AutoModelForCausalLM.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="memory-manager", torch_dtype="auto", device_map="auto")
35
- tokenizer = AutoTokenizer.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="memory-manager")
36
  ```
37
 
38
  ## Training
39
 
40
  - Base model: `Qwen/Qwen2.5-7B-Instruct`
41
- - Algorithm: LoGo-GRPO (turn-level + token-level advantage), curriculum-trained 8-session → 16-session → 32-session
42
- - Reward: per-session cumulative F1 against gold QA + a memory-compression penalty (λ=0.3)
43
  - Judge for reward/logging during training: GPT-OSS-120B
44
 
45
  ## Citation
 
24
 
25
  See the paper's `tab:different-answer-agent` for how this same memory manager performs when paired with other answer agents (untrained Qwen-7B, RL-trained Qwen-7B, etc.).
26
 
27
+ ## Answer agent (`answer-agent/`)
28
+
29
+ Also included is `sft_cont_step55`, a separately SFT+RL-trained Qwen2.5-7B-Instruct **answer agent**: given the memory manager's maintained memory store, it answers the held-out questions. Paired with the `memory-manager/` checkpoint above, this is the paper's highest-F1 configuration ("Memory-R2 (SFT-RL)" in `tab:main`):
30
+
31
+ | Memory manager | Answer agent | F1 | BLEU-1 | LLM-judge (gpt-4o-mini) |
32
+ | --- | --- | ---: | ---: | ---: |
33
+ | memory-manager/ | answer-agent/ | **51.46** | **44.84** | 69.03 |
34
+ | memory-manager/ | GPT-OSS-120B (untrained) | 49.29 | 43.64 | **86.08** |
35
+
36
  ## Usage
37
 
38
+ These models are two halves of a two-agent pipeline (memory manager + answer agent) and are not intended as general-purpose chat models on their own. Full inference code and the memory-store protocol are in the [project repository](https://github.com/) (see the paper for the official release).
39
 
40
  ```python
41
  from transformers import AutoModelForCausalLM, AutoTokenizer
42
 
43
+ memory_manager = AutoModelForCausalLM.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="memory-manager", torch_dtype="auto", device_map="auto")
44
+ answer_agent = AutoModelForCausalLM.from_pretrained("ahmedehabb/memory-r2-7b", subfolder="answer-agent", torch_dtype="auto", device_map="auto")
45
  ```
46
 
47
  ## Training
48
 
49
  - Base model: `Qwen/Qwen2.5-7B-Instruct`
50
+ - Memory manager: LoGo-GRPO (turn-level + token-level advantage), curriculum-trained 8-session → 16-session → 32-session; reward = per-session cumulative F1 against gold QA + a memory-compression penalty (λ=0.3)
51
+ - Answer agent: SFT warm-start followed by an RL continuation (answer-F1 reward) against the memory manager's rollouts
52
  - Judge for reward/logging during training: GPT-OSS-120B
53
 
54
  ## Citation