Manas-64M · Agentic RL

The SFT model trained with reinforcement learning on multi-turn tool use: up to 3 turns of <tool_call> → tool result → answer, rewarded when the final answer matches a verifiable ground truth.

Try it: chat Space · Every stage: collection · Code: github.com/kuluruvineeth/manas · Data: manas_dataset

Results

first 3 log points last 3 log points
mean reward -0.67 0.21
KL to the SFT reference 0.022 0.053
tool calls per trajectory 2.54 1.54
unfinished trajectories 75% 0%

800 steps in 87 minutes, one log point every 50 steps. Each point is a small batch, so read the trend, not single values.

training curves

Benchmarks

Zero-shot multiple choice, 500 items per task, scored by comparing answer log-likelihoods (scripts/evaluate.py).

benchmark accuracy random chance items
arc_challenge 21.0% 25.1% 500
arc_easy 32.4% 25.0% 500
hellaswag 35.0% 25.0% 500
mmlu 26.0% 25.0% 500
openbookqa 25.4% 25.0% 500
average 28.0%

A 64M model has far less capacity than these benchmarks assume, so scores sit near random chance. They are published because hiding them would be dishonest, and they move by about a point between stages, which is within noise at 500 items.

Use

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "kuluruvineeth/manas-64m-agent"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)

messages = [{"role": "user", "content": "What is the capital of France?"}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**inputs, max_new_tokens=60, do_sample=False)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

The export is a standard transformers checkpoint, so vLLM and SGLang load it with no custom code (vllm serve kuluruvineeth/manas-64m-agent).

Training

  • Lineage: starts from manas-64m-full-sft.
  • Method: GRPO loss over 4 rollouts per task, β = 0.1 KL to the SFT reference, up to 256 new tokens per turn.
  • Data: agent_rl.jsonl: tasks from 6 deterministic generators over mock tool tables, so every ground truth is correct by construction (sources and licenses).
  • Trainer: trainer/train_agent.py, peak learning rate 3e-7 (from the training log).
  • Hardware: one A100 or H100 GPU on Modal (gpu/modal_train.py).
  • Architecture: Qwen3 layout: 768 hidden, 8 layers, 8 query / 4 key-value heads (head dim 96), SwiGLU MLP of width 2,432, RoPE θ = 10⁶, tied embeddings, 6,400-token byte-level BPE vocabulary. 63,912,192 parameters.

Limitations

English only. At this size the model has little world knowledge: it is confidently wrong about facts that need it, loses the thread in long conversations, and should not be used for anything that matters. It is a teaching and research artifact for studying how each training stage changes a model.

Files

  • model.safetensors, config.json, tokenizer.json, chat_template.jinja: the transformers export
  • agent_768.pth: the raw training checkpoint for the training code
  • agent_768_metrics.jsonl, agent_768_curves.png: the training log and its plot
Downloads last month
10
Safetensors
Model size
63.9M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kuluruvineeth/manas-64m-agent

Finetuned
(5)
this model

Dataset used to train kuluruvineeth/manas-64m-agent

Space using kuluruvineeth/manas-64m-agent 1

Collection including kuluruvineeth/manas-64m-agent