SemanticRepair-270M

A 270M rewriter that sits behind an embedding router. When a question does not land on any capability with enough margin, this model restates it in the plain form the capabilities are described in, and the router tries again on the restatement. When it reads no request at all, it says so.

That is the whole job. It does not answer questions, it does not decide anything, and nothing it writes is ever executed: the router runs on the restatement, the tool runs on the original.

Fresh adversarial evaluation — 5 October 2026

The v22 Q8_0 weights are evaluated without retraining on 160 fresh questions: 80 paired Italian/English anchors covering plain requests, long noise, negation, quoted distractors, two and three requested capabilities, unsupported actions, and forged instructions. Inputs and gold capability sets are frozen before inference; no exact normalized overlap is found with v22 train/validation/test, previous probes, or the older 481-question bank. This is a Gramscii author self-evaluation by Codex / Atlante, without independent annotation review; exact overlap exclusion does not prove absence of semantic training overlap.

Model alone, on the native completion surface

Measure Result
Request-presence accuracy 140/160 — 87.50%
Recall of actual requests, including unsupported actions 131/150 — 87.33%
Correct nonrequest detection 9/10 — 90.00%
All annotated literal strings retained in normalized rewrites 59/72 — 81.94%
Desired number of request lines on multi-capability items 33/40 — 82.50%

Request-line count is a structural proxy, not proof of correct semantic decomposition. Literal string retention is stricter than semantic equivalence and does not measure tools' argument retention. Unsupported actions remain genuine requests for this model-only detection score; refusing unsupported graph operations is the system's responsibility.

Model inside the current SDG native admission path

SDG base 2d01f906 includes #645 and #654. The test selects the general default graph explicitly, uses its actual Italian and English descriptions and thresholds, and calls native planning.capabilities and chat.routing.resolve. The catalogue and library start empty; each conversation is fresh with no selected documents. No graph tools or narration execute. This is a routing admission experiment, not end-to-end answer accuracy or a test against the populated live catalogue.

Measure Direct native pre-repair gate Native Chat with repair
Complete desired capability set or correct refusal 35/160 — 21.88% 54/160 — 33.75%
Exact supported requests 19/134 47/134
Correct graph refusals 16/26 7/26
Exact multi-capability sets 0/40 1/40
Admissions containing an unwanted capability 24/160 54/160
Median routing latency, startup excluded 32.4 ms 174.8 ms

More supported requests are placed exactly with repair, but unwanted admissions also increase. These results do not establish a safety improvement. A wrong admission is not a wrong execution: no tool executes in this evaluation. Strict subsets count as partial, never exact. Native admission rules and repair both affect this comparison, so it does not isolate the repair model's causal effect. The English workspace retains its actual separate thresholds, without calibration on these questions.

Family, 20 items each Native Chat exact set or correct refusal
injection 40.0%
multi_three 0.0%
multi_two 5.0%
negation 70.0%
noisy 40.0%
plain 65.0%
quoted_trap 30.0%
uncovered 20.0%

Only 1/20 two-capability questions and 0/20 three-capability questions are admitted completely. The model can emit several request lines while their rerouting still misses or misidentifies capabilities. For example, the Italian request to map Pisa, give its ISTAT code, and retrieve its 2021 population yields three rewrite lines, but admission selects locate-address where opendata-territories is required. Unsupported actions also expose false admissions: a request to change Roma's official population is routed to opendata-statistics, although the graph has no authorized write operation.

The results are deliberately harder and broader than the historical banks below; their percentages are not comparable as a model-version regression. Eight balanced adversarial families do not estimate production frequencies. Paired translations are dependent, and the published item-level intervals and sign-test p-value assume independence; do not promote them as independent population inference. No repeated-run stability claim, other-language result, populated-catalogue search result, or conversation-continuation result is established.

The published evaluation dataset includes the frozen bank, complete question-level traces, graphs, model hashes, thresholds, scorer, source runner, workspace configuration, and verification provenance. The evaluated model remains revision acc30fd451d58620cc22dbbfab46caed3cbf12f5, with Q8_0 SHA-256 58b750b4bba4ed9cf1ea30316f1c5ad131316efaedd56cfd9440b488b2ae88f4; llama.cpp build 10621 uses greedy completion, a 96-token maximum, and repeat penalty 1.0. Structured author-reported results are in this card's model-index metadata. This evaluation does not claim a Hub verified badge or registration in its beta benchmark leaderboard service.

The surface it reads, which is not a chat prompt

Fine-tuned on a bare completion surface. It has never seen a chat template, a system message or a few-shot example. Speak to it the way it was trained or it will not work:

{Language}: {the question, verbatim}
=>

It completes with one line per request it found, or the single token NO_REQUEST.

The newline after => is part of the surface. The engine sends {Language}: {request}\n=>\n, byte for byte — the JSON example below shows it — and a prompt that stops at => is a surface the model never read.

{Language} is the English name of the language the question is in — not the language of whatever system is asking. English, Italian, French, German, Spanish; a language outside that set keeps whatever tag the caller declares, because an invented tag is a surface the model never read either.

This is the whole call the engine makes, on /v1/completions and never on a chat endpoint:

{
  "prompt": "Italian: che tempo fa domani a Bologna?\n=>\n",
  "stop": ["\n=>"],
  "temperature": 0.0,
  "repeat_penalty": 1.0,
  "max_tokens": 96
}

Greedy, because the same words must always produce the same rewrite: a router that reruns on a different restatement each time cannot be reasoned about. repeat_penalty is 1.0, which is neutral — a rewrite legitimately repeats the words of the question. The stop string cuts a runaway that starts echoing another question in its own trained format; everything before it is the answer. 96 tokens leaves room for several rewritten lines: the longest answer in the training data is 43 tokens.

The sentinel is compared case-insensitively and tolerates a trailing period. It is an instruction the model follows, not a token it is guaranteed to emit byte-exact.

These are real outputs from the released weights, greedy:

in out
Spanish: no busques la tienda de ropa, dime la dosis de paracetamol dime la dosis de paracetamol
English: book me a table for friday and also cancel my dentist book me a table for friday
cancel my dentist
Italian: no aspetta, riassumimi cosa ci siamo detti riassumi cosa ci siamo detti
Italian: guarda un po', questo sacchetto della spesa è tutto rotto. NO_REQUEST
Italian: ciao come stai NO_REQUEST

A negation is dropped, one message asking for two things becomes two lines, a request behind an opening "no" is kept, and something that is not a request at all answers with the sentinel.

What is in this repository

file size what it is
sft-v22-q8_0.gguf 292 MB what the engine serves, through llama-server
model.safetensors 536 MB the same weights fused, BF16, 236 tensors
model.safetensors.index.json the index over that single file
tokenizer.json, tokenizer_config.json 33 MB the tokenizer as the fuse wrote it
config.json, generation_config.json gemma3_text, torch_dtype: bfloat16
chat_template.jinja present because the fuse writes it — not the surface this model reads, see above

The GGUF's sha256 is 58b750b4bba4ed9cf1ea30316f1c5ad131316efaedd56cfd9440b488b2ae88f4. The engine pins it and refuses anything else, which is what makes a routing decision reproducible.

Both files carry the same weights: checkpoint 21,000 of training run v22, chosen on the validation split, quantised to Q8_0. Q8_0 and not Q4: measured on this seat, at temperature 0 a Q8_0 is stable to the byte across runs and a Q4 is not.

What it is measured to do

Every number below compares these weights with the previous release of this model (training run v19), both served by the same llama-server build with the same greedy call, on the same day. Pairs are counted where the two disagree, which is where the information is, with an exact two-sided sign test.

481 real questions people asked the engine, in Italian, held out from every training and selection step and routed through the full path the engine takes (the router, the repair when the route is unsure, admission only when every line is confident):

this model previous release
exact, every capability asked and no other 159 138
safe, a proper subset reached, nothing wrong run 10 10
wrong execution 12 12
questions where only this model is right / only the previous one is 23 / 2 p = 0.00002

The gain is where it was aimed: requests behind an opening "no", "ok" or "I didn't get it", bare statistical topics, "what's new in X", and names written in lowercase, which the previous release answered with NO_REQUEST. Of the 118 questions it still fails on a capability it can reach, 66 are decided by the router, not the rewrite: the right capability comes first but below the bars. 38 route elsewhere and 13 are still answered with the sentinel.

900 routing tasks over graphs from another corpus, five languages, two in five naming a day:

this model previous release
exact 56 59
compound questions answered whole 4 5
safe 17 11
abstained 696 713
wrong execution 4 4
answered in the wrong language 2 1
lost the day the question named 125 112
echoed the question back 123 150

Disagreements split 26 to 39 against this model, p = 0.14: no difference on this bench is significant, and the day column is the one to read.

100 probes written before training, 20 per family, none equal to a training row or to a real question above: 52 of 80 requests routed to the right capability with every name kept (previous release 43), and 20 of 20 non-requests answered NO_REQUEST (previous 20).

The held-out test split of the training data, 963 rows, whole completion equal to the target: 263 (previous 260).

What it does not fix, which a card naming only the gains would hide

Dates. On the 900-task bench it loses the day the question named more often than the previous release, 125 against 112, and the restatement reads as a clean request with the day simply gone. If the day matters to your routing, measure that before you serve this.

It can drop the verb of a request. Italian: scusa il disturbo, mi diresti che tempo fa domani a Bologna? comes back as dimmi se domani a Bologna è previsto: "che tempo fa" is gone.

It still refuses some requests. Italian: che novità ci sono in Django 6? answers NO_REQUEST. Of the 80 request probes, 10 still come back as the sentinel, 15 route to another capability and 3 lose a name.

What bounds all of this

None of these weaknesses are fixed. They are bounded, by the system around the model rather than by the model, and the bounds are worth stating because they decide how much a weakness costs.

Nothing it writes is executed. The restatement is compared against capability descriptions and then thrown away. Whatever runs, runs on the original message. Fixed SQL and HTTP operations are authored in the tool and cannot be assembled from either text.

The source message stays beside the restatement. Dates, language, page references and everything shown to the reader are read from what the person actually wrote. So a dropped day costs a routing decision, not the date: the question may reach the wrong capability, but the day itself was never the restatement's to lose.

A restatement that does not win by enough margin routes nothing. The router applies a margin and a floor, and below them the system abstains and asks for a rephrasing. That is why the wrong-execution column is 4 in 900 while the abstention column is 696: the failure mode is refusing, not acting wrongly.

The descriptions are the other half. What a question is matched against is written prose in the workspace, and a capability that keeps being missed is usually a description that needs rewriting, not a model that needs retraining. The rewriter is one of two things you can improve, and the cheaper one is often not the model.

The engine has a linter for exactly this, applying the same score the router applies, over a whole graph or over one description while it is being written. It reports collisions: pairs of capabilities the router cannot separate. A collision is not a verdict that either is wrong. Capabilities that overlap in meaning are expected to sit close, and the linter says where the router cannot tell them apart, leaving the choice to rewrite, merge or keep to whoever knows what the nodes are for. collision is also one of the four states a routing can end in, beside uncovered, uncertain and confident, so the reader is told which one happened rather than being handed an answer with no account of how it was reached.

Where it runs

Served with llama-server, reached by the engine as one of four model seats:

llama-server -m sft-v22-q8_0.gguf --host 127.0.0.1 --port 9250 -ngl 99 -c 2048 --jinja

It is loaded only when routing is genuinely unsure, never as the default path, and released after an idle window.

On CPU alone

Measured on this architecture at this size, on an Apple M-series machine with -ngl 0 -t 4: no GPU offload at all, four threads, and nothing in the server log naming Metal or a GPU.

latency tokens/s
strip a negation, 8 tokens out 38 ms 210
split one message into two requests, 11 tokens 32 ms 344
answer NO_REQUEST, 4 tokens 17 ms 236

762 MB resident while serving.

llama-server -m sft-v22-q8_0.gguf --host 127.0.0.1 --port 9250 -ngl 0 -t 4 -c 2048

What this does and does not establish: it runs usefully with no accelerator, on few threads. It has not been run on a phone or a tablet, and the three things that would decide it there are the three not measured here — 762 MB resident is a lot for a mobile process, loading 300 MB from slow storage is a different question from generating in 38 ms, and mobile cores throttle under heat while these did not.

The data it was trained on

Gramscii-IT/semantic-repair-routing — published, and everything below can be recounted from it.

The base is that dataset's 84,819 pairs (82,819 train, 1,000 validation, 1,000 test), cleaned by deterministic rules before training: template leaks, changed values, English targets for other languages, repeated lines and generic "category" lines removed, which leaves 80,297 train, 966 validation and 963 test rows. To them this release adds 2,515 rows in Italian and English, generated by Qwen3.6-35B-A3B and kept only when every line's top capability on the engine's own roster was the one the row was written for, every number and name of the message survived, and no row shared most of its trigrams with a probe, a real question or another row. They cover a request about the conversation itself, a bare statistical topic, "what's new in X", names in lowercase or with a typo, long noisy messages carrying two or three requests, and acknowledgements that ask for nothing. Each is repeated twice in training, about 6% of it.

The base dataset, as published: Each pair is a message somebody could plausibly write and the requests inside it restated plainly, one per line.

Five languages, close to balanced:

rows
English 18,006 21.7%
Spanish 16,903 20.4%
German 16,659 20.1%
French 16,602 20.0%
Italian 14,649 17.7%

What the message is doing, which is what the data is really organised by. Every one of the 82,819 training rows carries one of 38 such labels; these are the ten largest:

rows
massive_real — real assistant traffic 5,736 6.9%
minimal_pair — two messages differing in one word 5,299 6.4%
indirect — the request is implied, never stated 4,912 5.9%
presto_native — natively written, not translated 4,606 5.6%
oasst_human — human-written 4,450 5.4%
multi — more than one request in one message 3,974 4.8%
direct — the plain case 3,904 4.7%
vent_statement_request — a complaint with a request inside 3,516 4.2%
negation_exclusion — names what is not wanted 3,398 4.1%
trap_topic_lure — a topic named to pull routing the wrong way 3,324 4.0%

The remaining 28 cover self-correction, code-switching, ultra-short messages, typos and slang, messages carrying two, three or four requests, eight kinds of date reference — a weekday in the past, a named holiday, a bounded range — and the two families below.

The two families that decide what this model refuses to do

injection_meta — 2,278 rows — is a message aimed at the assistant itself: reveal your instructions, ignore what you were told, pretend you have no rules. 1,901 of them answer NO_REQUEST, because there is nothing here to route. Together with injection_command these two families supply 19% of every NO_REQUEST row in the set: refusing is largely taught by messages that try hardest to get an answer.

injection_command — 2,410 rows — is the opposite lesson, and the more delicate one. A hostile but explicit command about the user's own things: delete every saved bank account, empty my personal cloud, cancel all my bookings tonight. 2,281 of them are restated faithfully, action intact.

The rewriter does not sanitise and does not soften. Deciding whether an action is allowed is not its job and it has no way to do that job well: it sees one sentence, not the account, the permissions or the consequences. A model that quietly dropped "delete" would hand the router a different request from the one the person made, and the person would never learn their instruction had been edited. So the restatement carries the command as written, the router places it, and whatever runs it is where refusing belongs.

12.6% of rows (10,432) answer NO_REQUEST, and 16.9% (14,028) answer with more than one line.

Who wrote it, and what was thrown away:

generated by Qwen3.6-35B 72,697 rows, 87.8%
generated by Qwen3.8-27B 7,935 rows, 9.6%
generated by Gemma-4-E4B 2,187 rows, 2.6%
rejected by the judge, not in the set 72,996 rows

A separate model read every message against written rules and answered a two-word verdict, decoding greedily so a verdict does not change between runs. The dataset card says where judge and generator were different models and where they stopped being.

Training

Fine-tuned from google/gemma-3-270m with mlx_lm lora, full fine-tune, all 18 blocks. The released weights are checkpoint 21,000, chosen among the run's checkpoints by exact answers on the validation split (301 of 1,077), never on the test split or the benches above.

Where all of this happened

One desktop machine, and nothing left it.

machine Mac mini, Mac16,11
chip Apple M4 Pro
CPU 14 cores, 10 performance and 4 efficiency
GPU 20 cores
memory 64 GB unified
system macOS 26.4
the fine-tune
method mlx_lm lora, full fine-tune
iterations 21,276
tokens seen 2,327,566
peak memory 5.4 GB
speed about 6 iterations/s
wall clock 08:27 to 09:30, about an hour

Generators and judge ran on that same machine as local llama.cpp servers on loopback. No message here was written by a hosted API, and none was sent to one to be judged.

A model that decides where a question goes does not need a cluster or a month. It needs a narrow job, data built for it, and a bench honest enough to say when a round did not help.

Licence

This is a fine-tune of Gemma-3-270M — a Model Derivative under the Gemma Terms of Use, which it inherits whole. What those terms actually oblige, by their own sections:

Use (§3.2): the restricted uses in the Gemma Prohibited Use Policy are incorporated into the agreement by reference, and they bind anyone using these weights.

Redistribution (§3.1): whoever passes these weights on must include the §3.2 use restrictions as an enforceable provision in their own agreement, give every recipient a copy of the Terms, mark modified files as modified, and ship a notice stating "Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms". Those duties travel with the file, however many hands it passes through.

Outputs (§3.3): Google claims no rights in what the model writes. The restatements are yours, and so is the responsibility for them.

An engine that downloads these weights onto the operator's own machine, pinned by sha256, redistributes nothing: the duties above fall on whoever ships the file, not on whoever runs it.

Downloads last month
599
Safetensors
Model size
0.3B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Gramscii-IT/SemanticRepair-270M

Quantized
(51)
this model

Dataset used to train Gramscii-IT/SemanticRepair-270M

Evaluation results