File size: 1,231 Bytes
f6704da
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
# makeitwork1

Retriever500M — a 497M-parameter decoder-only transformer trained as a search agent.

## Architecture

Custom model defined in `src/model.py` (`Retriever500M`):
- d_model=1280, n_layers=23, n_heads=20, d_ff=3456
- RoPE positional embeddings
- Tied input/output embeddings

## Training

Two stages, both run on Modal (volume `retriever500m-data`):

1. **Pretraining** (1711 steps, corpus_curated.txt, lr=3e-4, seq_len=512):
   final EMA loss = 0.122
2. **SFT** (500 steps, sft_traces.jsonl + gold_traces.jsonl, lr=5e-5, seq_len=768):
   final EMA loss = 0.084

The checkpoint in this repo (`sft_latest.pt`) is the final SFT weights
(model_state_dict + config, step 500). Vocab size = 32009 (includes 9 special
agent tokens: system, user, assistant, search, result, evidence, reasoning,
finish, end).

## Loading

```python
import torch, sys
sys.path.insert(0, "src")
from model import ModelConfig, Retriever500M

ckpt = torch.load("sft_latest.pt", map_location="cuda", weights_only=False)
config = ModelConfig(**ckpt["config"])
model = Retriever500M(config).to("cuda")
model.load_state_dict(ckpt["model_state_dict"])
model.eval()
```

Tokenizer: `tokenizer/tokenizer_agent.json` (HuggingFace `tokenizers` library).