laya-context-prefilter

Laya multilingual (322M, mmBERT), fine-tuned to decide whether a Claude Code skill, rule, agent or MCP tool is useful for a user's request. It is the default model of claude-decide, a Claude Code plugin that passes Claude only the items a request needs.

This is v2, trained on a bilingual set (English and Spanish) with a larger catalog and more hard negatives. The first version stays available as the v1 revision.

Contract

  • State: {"request": "<the user's message>", "step": "<tool and main input, or empty>"}, for example {"request": "fix the flaky test", "step": "Bash npm test"}.
  • Question: one noul per candidate, with exactly this text:
Is this Claude Code {type} useful for handling the user's request? {type} «{name}»: {description}
  • Answer: noul is a calibrated probability, so a fixed threshold works (0.5) and no per-item prior is needed. Requests and descriptions can be in English, Spanish or Portuguese; the question stays in English.
import laya
from huggingface_hub import snapshot_download

agent = laya.load(snapshot_download("Zamax14/laya-context-prefilter"))  # revision="v1" for the first version
q = "Is this Claude Code skill useful for handling the user's request? skill «git-commit»: Writes conventional commit messages."
agent.predict({"request": "commit my changes", "step": ""}, {"git-commit": {"type": "noul", "instructions": q}})

Results

Held-out catalogs: pairs from repositories and MCP servers never seen in training, checked by two LLM judges.

Test Laya base v1 v2
English (3,120 pairs) 41% · Brier 0.403 90.3% · 0.078 90.3% · 0.074
Spanish (3,070 pairs, request and description translated) — 90.1% · 0.081 89.9% · 0.076

Ranking on a real catalog: 39 requests with hand-labelled answers against 49 installed items; three of the requests need nothing.

Model hit@1 hit@3 recall@5 MRR Items at 0.5 Precision at 0.5 Recall at 0.5
Laya base (with prior) 0.33 0.50 0.50 0.46 — — —
v1 0.86 1.00 0.94 0.93 2.8 0.41 0.82
v2 0.92 0.97 0.97 0.95 2.5 0.48 0.86
  • At 0.5, a request that needs nothing gets nothing, with both versions.

In claude-decide: on a project with 25 skills, 10 agents and 5 rules, 10 real tasks run 3 times without and with the plugin, with v2 of this model:

Claude model Tokens per task Tokens per turn Cost per task Tasks solved
Opus 5.5 134k → 81k (−39%) 25.0k → 17.2k (−31%) $0.209 → $0.132 (−37%) 30 → 30

With v1, Sonnet 5.5 used 35% fewer tokens per task and Haiku 4.5 14% fewer. The full comparison is in the claude-decide README.

  • Latency: about 0.24 s for 49 candidates on a laptop GPU (RTX 4050), and about 4.6 s on CPU.

Training

The model was trained with Laya-Finetune (context_prefilter task).

  1. Catalog: 3,351 skills, agents, rules and MCP tools, from public MIT and Apache repositories and the Docker MCP registry, plus synthetic MCP tools and rules. About 15% of the first round's sources were held out entirely.
  2. Requests: 15,000, written by qwen3.6:35b for items drawn from that catalog. They are in English, Spanish and Portuguese, with varied length and with and without a tool step.
  3. Negatives: each request is paired with the 5 most similar items (bge-m3 embeddings) and with 5 random ones.
  4. Judges: gemma4:31b checked every request and every near pair; qwen3.5:122b re-read what it rejected. When the two judges agreed against the constructed label, the pair was relabelled.
  5. Spanish: qwen3.6:35b translated every description and every request not in Spanish. Each pair got one Spanish copy with the same label: the description in Spanish, and half of the time the request too. A pair and its copy stay on the same side of the train/validation split.
  6. Training: 202,532 pairs (104k English, 98k Spanish), RLCD with calibration over the whole model, at 1,024 tokens of context. The best validation came after the first epoch.

Limitations

  • It judges an item from its description alone: an item with a vague description is harder to pick.
  • It was trained with 1k tokens of context. Longer states, such as the recent turns of a session, need more training.
  • About one relevant item in seven is missed at the 0.5 threshold.

License

Apache-2.0, inherited from Laya.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Zamax14/laya-context-prefilter

Finetuned
(34)
this model