Makalu

Makalu is a 0.6B embedding model trained to make fast, calibrated decisions. It scores a situation against a list of options in one pass, and it estimates whether a finished coding-agent attempt will pass. It is Qwen3-Embedding-0.6B with a LoRA adapter (rank 16, all 7 projections in all 28 layers) merged into its weights, plus two small scoring rules.

use input output
retrieval a question and passages cosine similarity
choices a question or agent state and its options one probability per option
verification a summary of a coding attempt P(attempt passes)

Architecture

Text is encoded with left padding and pooled at its last token, then L2-normalised to a 1024-d unit vector. The tokenizer in this repo appends <|im_end|><|endoftext|> to every input, exactly as in training.

  • Choice scorer: p = softmax(α · cos(state, optionₖ)), with α = 14.15 (in makalu_config.json). The state is Instruct: <instruction>\nQuery: <question>; each option is Question: <question>\nAnswer: <option>.
  • Verifier head: P(pass) = σ(w · e + b), a linear head on the summary embedding e (makalu_heads.safetensors, 1,025 parameters).

Usage

With the helper in this repo (torch and transformers only):

from huggingface_hub import hf_hub_download
import importlib.util, sys
spec = importlib.util.spec_from_file_location("makalu", hf_hub_download("spandyie/makalu", "makalu.py"))
makalu = importlib.util.module_from_spec(spec); spec.loader.exec_module(makalu)

m = makalu.Makalu.from_pretrained("spandyie/makalu")          # downloads the repo; GPU if available

m.retrieve(["who wrote hamlet"], ["Hamlet is a tragedy by William Shakespeare.", "Paris is in France."])
m.score_choices("Which gas do plants absorb for photosynthesis?", ["oxygen", "carbon dioxide", "helium", "neon"])

s = m.make_summary(task="<issue text>", diff="<final git diff>",
                   last_test="<output of the last test command>", last_actions="<last few commands>")
m.p_pass([s])                                                   # tensor([P(pass)])

With sentence-transformers (embeddings only; the query prompt adds the retrieval instruction):

from sentence_transformers import SentenceTransformer
model = SentenceTransformer("spandyie/makalu")
q = model.encode(["who wrote hamlet"], prompt_name="query")
d = model.encode(["Hamlet is a tragedy by William Shakespeare."])
print(model.similarity(q, d))

With plain transformers:

import torch, torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("spandyie/makalu", padding_side="left", truncation_side="left")
enc = AutoModel.from_pretrained("spandyie/makalu").eval()
bt = tok(["Instruct: Given a question, retrieve the passage that answers it.\nQuery: who wrote hamlet"],
         padding=True, truncation=True, max_length=256, return_tensors="pt")
with torch.no_grad():
    e = F.normalize(enc(**bt).last_hidden_state[:, -1], dim=-1)

Use max_length=256 for questions, options and passages, and 2048 for attempt summaries, as in training.

Training

The adapter was trained in four stages on one GPU, about 39 GPU hours in total. Each stage started from the previous one. The base weights stayed frozen.

stage data objective steps
retrieval 100k PAQ + StackExchange question–passage pairs InfoNCE, 1 hard negative per query 1,562
multiple choice 105,440 OpenBookQA + MMLU questions Brier loss on softmax(α·cos) 4,818
agent actions 306,965 decisions from 23 agent datasets (neulab/agent-data-collection) GRPO, reward y − (q − y)², KL 0.02 3,120
verification 20,273 graded attempts on 2,154 tasks from 32 agents Brier on P(pass) + hard-negative pair ranking + 40% replay of the earlier two stages 4,059

In the verification stage, each attempt's logit got a fixed offset equal to its agent's training pass-rate log-odds, so the head learns from the content of the attempt and not from which agent wrote it. The released head scores content only.

Results

Held-out results for this checkpoint, against plain Qwen3-Embedding-0.6B:

benchmark metric Qwen3-Embedding-0.6B Makalu
retrieval, 2,000 questions vs 21,600 passages top-1 0.770 0.650
multiple choice (OpenBookQA + MMLU val) accuracy 0.283 0.587
Brier ↓ 0.751 0.542
agent next action (21 datasets, val, 4 options) accuracy 0.400 0.714
SWE-bench Verified (95 tasks, test) best-of-4 pass rate 0.664 0.732
AUROC 0.59 0.739
calibration error ↓ — 0.050
Terminal-Bench 2.1 (89 tasks, test, unseen) best-of-2 pass rate — 0.618

On SWE-bench Verified, a random pick passes 0.602 of the time and a perfect pick 0.874. A baseline that ranks attempts only by how often their agent passes reaches 0.741, so Makalu matches it (95% intervals overlap) while reading only the attempt. Among attempts from the same agent, where that baseline is at 0.5, Makalu's AUROC is 0.672.

Limitations

  • Retrieval accuracy is lower than the base model, because the later stages did not train on retrieval. Use the base model if you only need retrieval.
  • Most numbers come from single training runs, and the verification sets are small (81 dev and 95 test tasks), so differences under about 0.01 are within noise.
  • Terminal-Bench has 2 attempts per task, which is a weak test of selection, and the verifier's calibration does not transfer to it (calibration error 0.22).
  • The verifier expects the summary format above. Other inputs will still embed, but its probabilities are only meaningful for that format.

License

Apache 2.0, the same as the base model. The training datasets carry their own licenses.

Downloads last month
8
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for spandyie/makalu

Adapter
(30)
this model

Datasets used to train spandyie/makalu