Instructions to use spandyie/makalu with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use spandyie/makalu with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="spandyie/makalu")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("spandyie/makalu") model = AutoModel.from_pretrained("spandyie/makalu", device_map="auto") - sentence-transformers
How to use spandyie/makalu with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("spandyie/makalu") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Makalu
Makalu is a 0.6B embedding model trained to make fast, calibrated decisions. It scores a situation against a list of options in one pass, and it estimates whether a finished coding-agent attempt will pass. It is Qwen3-Embedding-0.6B with a LoRA adapter (rank 16, all 7 projections in all 28 layers) merged into its weights, plus two small scoring rules.
| use | input | output |
|---|---|---|
| retrieval | a question and passages | cosine similarity |
| choices | a question or agent state and its options | one probability per option |
| verification | a summary of a coding attempt | P(attempt passes) |
Architecture
Text is encoded with left padding and pooled at its last token, then L2-normalised to a 1024-d unit vector. The tokenizer in this repo appends <|im_end|><|endoftext|> to every input, exactly as in training.
- Choice scorer: p = softmax(α · cos(state, optionₖ)), with α = 14.15 (in
makalu_config.json). The state isInstruct: <instruction>\nQuery: <question>; each option isQuestion: <question>\nAnswer: <option>. - Verifier head: P(pass) = σ(w · e + b), a linear head on the summary embedding e (
makalu_heads.safetensors, 1,025 parameters).
Usage
With the helper in this repo (torch and transformers only):
from huggingface_hub import hf_hub_download
import importlib.util, sys
spec = importlib.util.spec_from_file_location("makalu", hf_hub_download("spandyie/makalu", "makalu.py"))
makalu = importlib.util.module_from_spec(spec); spec.loader.exec_module(makalu)
m = makalu.Makalu.from_pretrained("spandyie/makalu") # downloads the repo; GPU if available
m.retrieve(["who wrote hamlet"], ["Hamlet is a tragedy by William Shakespeare.", "Paris is in France."])
m.score_choices("Which gas do plants absorb for photosynthesis?", ["oxygen", "carbon dioxide", "helium", "neon"])
s = m.make_summary(task="<issue text>", diff="<final git diff>",
last_test="<output of the last test command>", last_actions="<last few commands>")
m.p_pass([s]) # tensor([P(pass)])
With sentence-transformers (embeddings only; the query prompt adds the retrieval instruction):
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("spandyie/makalu")
q = model.encode(["who wrote hamlet"], prompt_name="query")
d = model.encode(["Hamlet is a tragedy by William Shakespeare."])
print(model.similarity(q, d))
With plain transformers:
import torch, torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("spandyie/makalu", padding_side="left", truncation_side="left")
enc = AutoModel.from_pretrained("spandyie/makalu").eval()
bt = tok(["Instruct: Given a question, retrieve the passage that answers it.\nQuery: who wrote hamlet"],
padding=True, truncation=True, max_length=256, return_tensors="pt")
with torch.no_grad():
e = F.normalize(enc(**bt).last_hidden_state[:, -1], dim=-1)
Use max_length=256 for questions, options and passages, and 2048 for attempt summaries, as in training.
Training
The adapter was trained in four stages on one GPU, about 39 GPU hours in total. Each stage started from the previous one. The base weights stayed frozen.
| stage | data | objective | steps |
|---|---|---|---|
| retrieval | 100k PAQ + StackExchange question–passage pairs | InfoNCE, 1 hard negative per query | 1,562 |
| multiple choice | 105,440 OpenBookQA + MMLU questions | Brier loss on softmax(α·cos) | 4,818 |
| agent actions | 306,965 decisions from 23 agent datasets (neulab/agent-data-collection) | GRPO, reward y − (q − y)², KL 0.02 | 3,120 |
| verification | 20,273 graded attempts on 2,154 tasks from 32 agents | Brier on P(pass) + hard-negative pair ranking + 40% replay of the earlier two stages | 4,059 |
In the verification stage, each attempt's logit got a fixed offset equal to its agent's training pass-rate log-odds, so the head learns from the content of the attempt and not from which agent wrote it. The released head scores content only.
Results
Held-out results for this checkpoint, against plain Qwen3-Embedding-0.6B:
| benchmark | metric | Qwen3-Embedding-0.6B | Makalu |
|---|---|---|---|
| retrieval, 2,000 questions vs 21,600 passages | top-1 | 0.770 | 0.650 |
| multiple choice (OpenBookQA + MMLU val) | accuracy | 0.283 | 0.587 |
| Brier ↓ | 0.751 | 0.542 | |
| agent next action (21 datasets, val, 4 options) | accuracy | 0.400 | 0.714 |
| SWE-bench Verified (95 tasks, test) | best-of-4 pass rate | 0.664 | 0.732 |
| AUROC | 0.59 | 0.739 | |
| calibration error ↓ | — | 0.050 | |
| Terminal-Bench 2.1 (89 tasks, test, unseen) | best-of-2 pass rate | — | 0.618 |
On SWE-bench Verified, a random pick passes 0.602 of the time and a perfect pick 0.874. A baseline that ranks attempts only by how often their agent passes reaches 0.741, so Makalu matches it (95% intervals overlap) while reading only the attempt. Among attempts from the same agent, where that baseline is at 0.5, Makalu's AUROC is 0.672.
Limitations
- Retrieval accuracy is lower than the base model, because the later stages did not train on retrieval. Use the base model if you only need retrieval.
- Most numbers come from single training runs, and the verification sets are small (81 dev and 95 test tasks), so differences under about 0.01 are within noise.
- Terminal-Bench has 2 attempts per task, which is a weak test of selection, and the verifier's calibration does not transfer to it (calibration error 0.22).
- The verifier expects the summary format above. Other inputs will still embed, but its probabilities are only meaningful for that format.
License
Apache 2.0, the same as the base model. The training datasets carry their own licenses.
- Downloads last month
- 8