|
Download README.md from Reizxn/makeitwork1: direct link, hf CLI and curl.
- Browser
- Download file 1.23 kB
-
https://huggingface.co/Reizxn/makeitwork1/resolve/main/README.md
- Command line
-
hf download hf://Reizxn/makeitwork1/README.md
-
curl -L -o README.md https://huggingface.co/Reizxn/makeitwork1/resolve/main/README.md
1.23 kB
makeitwork1
Retriever500M — a 497M-parameter decoder-only transformer trained as a search agent.
Architecture
Custom model defined in src/model.py (Retriever500M):
- d_model=1280, n_layers=23, n_heads=20, d_ff=3456
- RoPE positional embeddings
- Tied input/output embeddings
Training
Two stages, both run on Modal (volume retriever500m-data):
- Pretraining (1711 steps, corpus_curated.txt, lr=3e-4, seq_len=512): final EMA loss = 0.122
- SFT (500 steps, sft_traces.jsonl + gold_traces.jsonl, lr=5e-5, seq_len=768): final EMA loss = 0.084
The checkpoint in this repo (sft_latest.pt) is the final SFT weights
(model_state_dict + config, step 500). Vocab size = 32009 (includes 9 special
agent tokens: system, user, assistant, search, result, evidence, reasoning,
finish, end).
Loading
import torch, sys
sys.path.insert(0, "src")
from model import ModelConfig, Retriever500M
ckpt = torch.load("sft_latest.pt", map_location="cuda", weights_only=False)
config = ModelConfig(**ckpt["config"])
model = Retriever500M(config).to("cuda")
model.load_state_dict(ckpt["model_state_dict"])
model.eval()
Tokenizer: tokenizer/tokenizer_agent.json (HuggingFace tokenizers library).