laya-context-prefilter
Laya multilingual (322M, mmBERT), fine-tuned to decide whether a Claude Code skill, rule, agent or MCP tool is useful for a user's request. It is the default model of claude-decide, a Claude Code plugin that passes Claude only the items a request needs.
This is v2, trained on a bilingual set (English and Spanish) with a larger catalog and more hard negatives. The
first version stays available as the v1 revision.
Contract
- State:
{"request": "<the user's message>", "step": "<tool and main input, or empty>"}, for example{"request": "fix the flaky test", "step": "Bash npm test"}. - Question: one
noulper candidate, with exactly this text:
Is this Claude Code {type} useful for handling the user's request? {type} «{name}»: {description}
- Answer:
noulis a calibrated probability, so a fixed threshold works (0.5) and no per-item prior is needed. Requests and descriptions can be in English, Spanish or Portuguese; the question stays in English.
import laya
from huggingface_hub import snapshot_download
agent = laya.load(snapshot_download("Zamax14/laya-context-prefilter")) # revision="v1" for the first version
q = "Is this Claude Code skill useful for handling the user's request? skill «git-commit»: Writes conventional commit messages."
agent.predict({"request": "commit my changes", "step": ""}, {"git-commit": {"type": "noul", "instructions": q}})
Results
Held-out catalogs: pairs from repositories and MCP servers never seen in training, checked by two LLM judges.
| Test | Laya base | v1 | v2 |
|---|---|---|---|
| English (3,120 pairs) | 41% · Brier 0.403 | 90.3% · 0.078 | 90.3% · 0.074 |
| Spanish (3,070 pairs, request and description translated) | — | 90.1% · 0.081 | 89.9% · 0.076 |
Ranking on a real catalog: 39 requests with hand-labelled answers against 49 installed items; three of the requests need nothing.
| Model | hit@1 | hit@3 | recall@5 | MRR | Items at 0.5 | Precision at 0.5 | Recall at 0.5 |
|---|---|---|---|---|---|---|---|
| Laya base (with prior) | 0.33 | 0.50 | 0.50 | 0.46 | — | — | — |
| v1 | 0.86 | 1.00 | 0.94 | 0.93 | 2.8 | 0.41 | 0.82 |
| v2 | 0.92 | 0.97 | 0.97 | 0.95 | 2.5 | 0.48 | 0.86 |
- At 0.5, a request that needs nothing gets nothing, with both versions.
In claude-decide: on a project with 25 skills, 10 agents and 5 rules, 10 real tasks run 3 times without and with the plugin, with v2 of this model:
| Claude model | Tokens per task | Tokens per turn | Cost per task | Tasks solved |
|---|---|---|---|---|
| Opus 5.5 | 134k → 81k (−39%) | 25.0k → 17.2k (−31%) | $0.209 → $0.132 (−37%) | 30 → 30 |
With v1, Sonnet 5.5 used 35% fewer tokens per task and Haiku 4.5 14% fewer. The full comparison is in the claude-decide README.
- Latency: about 0.24 s for 49 candidates on a laptop GPU (RTX 4050), and about 4.6 s on CPU.
Training
The model was trained with Laya-Finetune (context_prefilter task).
- Catalog: 3,351 skills, agents, rules and MCP tools, from public MIT and Apache repositories and the Docker MCP registry, plus synthetic MCP tools and rules. About 15% of the first round's sources were held out entirely.
- Requests: 15,000, written by qwen3.6:35b for items drawn from that catalog. They are in English, Spanish and Portuguese, with varied length and with and without a tool step.
- Negatives: each request is paired with the 5 most similar items (bge-m3 embeddings) and with 5 random ones.
- Judges: gemma4:31b checked every request and every near pair; qwen3.5:122b re-read what it rejected. When the two judges agreed against the constructed label, the pair was relabelled.
- Spanish: qwen3.6:35b translated every description and every request not in Spanish. Each pair got one Spanish copy with the same label: the description in Spanish, and half of the time the request too. A pair and its copy stay on the same side of the train/validation split.
- Training: 202,532 pairs (104k English, 98k Spanish), RLCD with calibration over the whole model, at 1,024 tokens of context. The best validation came after the first epoch.
Limitations
- It judges an item from its description alone: an item with a vague description is harder to pick.
- It was trained with 1k tokens of context. Longer states, such as the recent turns of a session, need more training.
- About one relevant item in seven is missed at the 0.5 threshold.
License
Apache-2.0, inherited from Laya.
Model tree for Zamax14/laya-context-prefilter
Base model
convaiinnovations/laya-multilingual