# makeitwork1 Retriever500M — a 497M-parameter decoder-only transformer trained as a search agent. ## Architecture Custom model defined in `src/model.py` (`Retriever500M`): - d_model=1280, n_layers=23, n_heads=20, d_ff=3456 - RoPE positional embeddings - Tied input/output embeddings ## Training Two stages, both run on Modal (volume `retriever500m-data`): 1. **Pretraining** (1711 steps, corpus_curated.txt, lr=3e-4, seq_len=512): final EMA loss = 0.122 2. **SFT** (500 steps, sft_traces.jsonl + gold_traces.jsonl, lr=5e-5, seq_len=768): final EMA loss = 0.084 The checkpoint in this repo (`sft_latest.pt`) is the final SFT weights (model_state_dict + config, step 500). Vocab size = 32009 (includes 9 special agent tokens: system, user, assistant, search, result, evidence, reasoning, finish, end). ## Loading ```python import torch, sys sys.path.insert(0, "src") from model import ModelConfig, Retriever500M ckpt = torch.load("sft_latest.pt", map_location="cuda", weights_only=False) config = ModelConfig(**ckpt["config"]) model = Retriever500M(config).to("cuda") model.load_state_dict(ckpt["model_state_dict"]) model.eval() ``` Tokenizer: `tokenizer/tokenizer_agent.json` (HuggingFace `tokenizers` library).