Instructions to use Gramscii-IT/SemanticRepair-270M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Gramscii-IT/SemanticRepair-270M with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Gramscii-IT/SemanticRepair-270M") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Gramscii-IT/SemanticRepair-270M with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Gramscii-IT/SemanticRepair-270M:Q8_0 # Run inference directly in the terminal: llama cli -hf Gramscii-IT/SemanticRepair-270M:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Gramscii-IT/SemanticRepair-270M:Q8_0 # Run inference directly in the terminal: llama cli -hf Gramscii-IT/SemanticRepair-270M:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Gramscii-IT/SemanticRepair-270M:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Gramscii-IT/SemanticRepair-270M:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Gramscii-IT/SemanticRepair-270M:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Gramscii-IT/SemanticRepair-270M:Q8_0
Use Docker
docker model run hf.co/Gramscii-IT/SemanticRepair-270M:Q8_0
- LM Studio
- Jan
- vLLM
How to use Gramscii-IT/SemanticRepair-270M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Gramscii-IT/SemanticRepair-270M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Gramscii-IT/SemanticRepair-270M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Gramscii-IT/SemanticRepair-270M:Q8_0
- Ollama
How to use Gramscii-IT/SemanticRepair-270M with Ollama:
ollama run hf.co/Gramscii-IT/SemanticRepair-270M:Q8_0
- Unsloth Desktop
- MLX LM
How to use Gramscii-IT/SemanticRepair-270M with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Gramscii-IT/SemanticRepair-270M"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Gramscii-IT/SemanticRepair-270M" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Gramscii-IT/SemanticRepair-270M", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use Gramscii-IT/SemanticRepair-270M with Docker Model Runner:
docker model run hf.co/Gramscii-IT/SemanticRepair-270M:Q8_0
- Lemonade
How to use Gramscii-IT/SemanticRepair-270M with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Gramscii-IT/SemanticRepair-270M:Q8_0
Run and chat with the model
lemonade run user.SemanticRepair-270M-Q8_0
List all available models
lemonade list
- Atomic Chat
- SemanticRepair-270M
- Fresh adversarial evaluation — 5 October 2026
- The surface it reads, which is not a chat prompt
- What is in this repository
- What it is measured to do
- What it does not fix, which a card naming only the gains would hide
- Where it runs
- The data it was trained on
- Training
- Where all of this happened
- Licence
SemanticRepair-270M
A 270M rewriter that sits behind an embedding router. When a question does not land on any capability with enough margin, this model restates it in the plain form the capabilities are described in, and the router tries again on the restatement. When it reads no request at all, it says so.
That is the whole job. It does not answer questions, it does not decide anything, and nothing it writes is ever executed: the router runs on the restatement, the tool runs on the original.
Fresh adversarial evaluation — 5 October 2026
The v22 Q8_0 weights are evaluated without retraining on 160 fresh questions: 80 paired Italian/English anchors covering plain requests, long noise, negation, quoted distractors, two and three requested capabilities, unsupported actions, and forged instructions. Inputs and gold capability sets are frozen before inference; no exact normalized overlap is found with v22 train/validation/test, previous probes, or the older 481-question bank. This is a Gramscii author self-evaluation by Codex / Atlante, without independent annotation review; exact overlap exclusion does not prove absence of semantic training overlap.
Model alone, on the native completion surface
| Measure | Result |
|---|---|
| Request-presence accuracy | 140/160 — 87.50% |
| Recall of actual requests, including unsupported actions | 131/150 — 87.33% |
| Correct nonrequest detection | 9/10 — 90.00% |
| All annotated literal strings retained in normalized rewrites | 59/72 — 81.94% |
| Desired number of request lines on multi-capability items | 33/40 — 82.50% |
Request-line count is a structural proxy, not proof of correct semantic decomposition. Literal string retention is stricter than semantic equivalence and does not measure tools' argument retention. Unsupported actions remain genuine requests for this model-only detection score; refusing unsupported graph operations is the system's responsibility.
Model inside the current SDG native admission path
SDG base 2d01f906 includes #645 and #654. The test selects the general default graph explicitly, uses its actual Italian and English descriptions and thresholds, and calls native planning.capabilities and chat.routing.resolve. The catalogue and library start empty; each conversation is fresh with no selected documents. No graph tools or narration execute. This is a routing admission experiment, not end-to-end answer accuracy or a test against the populated live catalogue.
| Measure | Direct native pre-repair gate | Native Chat with repair |
|---|---|---|
| Complete desired capability set or correct refusal | 35/160 — 21.88% | 54/160 — 33.75% |
| Exact supported requests | 19/134 | 47/134 |
| Correct graph refusals | 16/26 | 7/26 |
| Exact multi-capability sets | 0/40 | 1/40 |
| Admissions containing an unwanted capability | 24/160 | 54/160 |
| Median routing latency, startup excluded | 32.4 ms | 174.8 ms |
More supported requests are placed exactly with repair, but unwanted admissions also increase. These results do not establish a safety improvement. A wrong admission is not a wrong execution: no tool executes in this evaluation. Strict subsets count as partial, never exact. Native admission rules and repair both affect this comparison, so it does not isolate the repair model's causal effect. The English workspace retains its actual separate thresholds, without calibration on these questions.
| Family, 20 items each | Native Chat exact set or correct refusal |
|---|---|
| injection | 40.0% |
| multi_three | 0.0% |
| multi_two | 5.0% |
| negation | 70.0% |
| noisy | 40.0% |
| plain | 65.0% |
| quoted_trap | 30.0% |
| uncovered | 20.0% |
Only 1/20 two-capability questions and 0/20 three-capability questions are admitted completely. The model can emit several request lines while their rerouting still misses or misidentifies capabilities. For example, the Italian request to map Pisa, give its ISTAT code, and retrieve its 2021 population yields three rewrite lines, but admission selects locate-address where opendata-territories is required. Unsupported actions also expose false admissions: a request to change Roma's official population is routed to opendata-statistics, although the graph has no authorized write operation.
The results are deliberately harder and broader than the historical banks below; their percentages are not comparable as a model-version regression. Eight balanced adversarial families do not estimate production frequencies. Paired translations are dependent, and the published item-level intervals and sign-test p-value assume independence; do not promote them as independent population inference. No repeated-run stability claim, other-language result, populated-catalogue search result, or conversation-continuation result is established.
The published evaluation dataset includes the frozen bank, complete question-level traces, graphs, model hashes, thresholds, scorer, source runner, workspace configuration, and verification provenance. The evaluated model remains revision acc30fd451d58620cc22dbbfab46caed3cbf12f5, with Q8_0 SHA-256 58b750b4bba4ed9cf1ea30316f1c5ad131316efaedd56cfd9440b488b2ae88f4; llama.cpp build 10621 uses greedy completion, a 96-token maximum, and repeat penalty 1.0. Structured author-reported results are in this card's model-index metadata. This evaluation does not claim a Hub verified badge or registration in its beta benchmark leaderboard service.
The surface it reads, which is not a chat prompt
Fine-tuned on a bare completion surface. It has never seen a chat template, a system message or a few-shot example. Speak to it the way it was trained or it will not work:
{Language}: {the question, verbatim}
=>
It completes with one line per request it found, or the single token
NO_REQUEST.
The newline after => is part of the surface. The engine sends
{Language}: {request}\n=>\n, byte for byte — the JSON example below
shows it — and a prompt that stops at => is a surface the model never
read.
{Language} is the English name of the language the question is in —
not the language of whatever system is asking. English, Italian,
French, German, Spanish; a language outside that set keeps whatever
tag the caller declares, because an invented tag is a surface the model
never read either.
This is the whole call the engine makes, on /v1/completions and never on
a chat endpoint:
{
"prompt": "Italian: che tempo fa domani a Bologna?\n=>\n",
"stop": ["\n=>"],
"temperature": 0.0,
"repeat_penalty": 1.0,
"max_tokens": 96
}
Greedy, because the same words must always produce the same rewrite: a
router that reruns on a different restatement each time cannot be reasoned
about. repeat_penalty is 1.0, which is neutral — a rewrite legitimately
repeats the words of the question. The stop string cuts a runaway that
starts echoing another question in its own trained format; everything
before it is the answer. 96 tokens leaves room for several rewritten lines:
the longest answer in the training data is 43 tokens.
The sentinel is compared case-insensitively and tolerates a trailing period. It is an instruction the model follows, not a token it is guaranteed to emit byte-exact.
These are real outputs from the released weights, greedy:
| in | out |
|---|---|
Spanish: no busques la tienda de ropa, dime la dosis de paracetamol |
dime la dosis de paracetamol |
English: book me a table for friday and also cancel my dentist |
book me a table for fridaycancel my dentist |
Italian: no aspetta, riassumimi cosa ci siamo detti |
riassumi cosa ci siamo detti |
Italian: guarda un po', questo sacchetto della spesa è tutto rotto. |
NO_REQUEST |
Italian: ciao come stai |
NO_REQUEST |
A negation is dropped, one message asking for two things becomes two lines, a request behind an opening "no" is kept, and something that is not a request at all answers with the sentinel.
What is in this repository
| file | size | what it is |
|---|---|---|
sft-v22-q8_0.gguf |
292 MB | what the engine serves, through llama-server |
model.safetensors |
536 MB | the same weights fused, BF16, 236 tensors |
model.safetensors.index.json |
the index over that single file | |
tokenizer.json, tokenizer_config.json |
33 MB | the tokenizer as the fuse wrote it |
config.json, generation_config.json |
gemma3_text, torch_dtype: bfloat16 |
|
chat_template.jinja |
present because the fuse writes it — not the surface this model reads, see above |
The GGUF's sha256 is
58b750b4bba4ed9cf1ea30316f1c5ad131316efaedd56cfd9440b488b2ae88f4.
The engine pins it and refuses anything else, which is what makes a routing
decision reproducible.
Both files carry the same weights: checkpoint 21,000 of training run v22, chosen on the validation split, quantised to Q8_0. Q8_0 and not Q4: measured on this seat, at temperature 0 a Q8_0 is stable to the byte across runs and a Q4 is not.
What it is measured to do
Every number below compares these weights with the previous release of this model (training run v19), both served by the same llama-server build with the same greedy call, on the same day. Pairs are counted where the two disagree, which is where the information is, with an exact two-sided sign test.
481 real questions people asked the engine, in Italian, held out from every training and selection step and routed through the full path the engine takes (the router, the repair when the route is unsure, admission only when every line is confident):
| this model | previous release | |
|---|---|---|
| exact, every capability asked and no other | 159 | 138 |
| safe, a proper subset reached, nothing wrong run | 10 | 10 |
| wrong execution | 12 | 12 |
| questions where only this model is right / only the previous one is | 23 / 2 | p = 0.00002 |
The gain is where it was aimed: requests behind an opening "no", "ok" or
"I didn't get it", bare statistical topics, "what's new in X", and names
written in lowercase, which the previous release answered with
NO_REQUEST. Of the 118 questions it still fails on a capability it can
reach, 66 are decided by the router, not the rewrite: the right capability
comes first but below the bars. 38 route elsewhere and 13 are still
answered with the sentinel.
900 routing tasks over graphs from another corpus, five languages, two in five naming a day:
| this model | previous release | |
|---|---|---|
| exact | 56 | 59 |
| compound questions answered whole | 4 | 5 |
| safe | 17 | 11 |
| abstained | 696 | 713 |
| wrong execution | 4 | 4 |
| answered in the wrong language | 2 | 1 |
| lost the day the question named | 125 | 112 |
| echoed the question back | 123 | 150 |
Disagreements split 26 to 39 against this model, p = 0.14: no difference on this bench is significant, and the day column is the one to read.
100 probes written before training, 20 per family, none equal to a
training row or to a real question above: 52 of 80 requests routed to the
right capability with every name kept (previous release 43), and 20 of 20
non-requests answered NO_REQUEST (previous 20).
The held-out test split of the training data, 963 rows, whole completion equal to the target: 263 (previous 260).
What it does not fix, which a card naming only the gains would hide
Dates. On the 900-task bench it loses the day the question named more often than the previous release, 125 against 112, and the restatement reads as a clean request with the day simply gone. If the day matters to your routing, measure that before you serve this.
It can drop the verb of a request. Italian: scusa il disturbo, mi diresti che tempo fa domani a Bologna? comes back as dimmi se domani a Bologna è previsto: "che tempo fa" is gone.
It still refuses some requests. Italian: che novità ci sono in Django 6? answers NO_REQUEST. Of the 80 request probes, 10 still come back as
the sentinel, 15 route to another capability and 3 lose a name.
What bounds all of this
None of these weaknesses are fixed. They are bounded, by the system around the model rather than by the model, and the bounds are worth stating because they decide how much a weakness costs.
Nothing it writes is executed. The restatement is compared against capability descriptions and then thrown away. Whatever runs, runs on the original message. Fixed SQL and HTTP operations are authored in the tool and cannot be assembled from either text.
The source message stays beside the restatement. Dates, language, page references and everything shown to the reader are read from what the person actually wrote. So a dropped day costs a routing decision, not the date: the question may reach the wrong capability, but the day itself was never the restatement's to lose.
A restatement that does not win by enough margin routes nothing. The router applies a margin and a floor, and below them the system abstains and asks for a rephrasing. That is why the wrong-execution column is 4 in 900 while the abstention column is 696: the failure mode is refusing, not acting wrongly.
The descriptions are the other half. What a question is matched against is written prose in the workspace, and a capability that keeps being missed is usually a description that needs rewriting, not a model that needs retraining. The rewriter is one of two things you can improve, and the cheaper one is often not the model.
The engine has a linter for exactly this, applying the same score the router
applies, over a whole graph or over one description while it is being
written. It reports collisions: pairs of capabilities the router cannot
separate. A collision is not a verdict that either is wrong. Capabilities
that overlap in meaning are expected to sit close, and the linter says where
the router cannot tell them apart, leaving the choice to rewrite, merge or
keep to whoever knows what the nodes are for. collision is also one of the
four states a routing can end in, beside uncovered, uncertain and
confident, so the reader is told which one happened rather than being
handed an answer with no account of how it was reached.
Where it runs
Served with llama-server, reached by the engine as one of four model seats:
llama-server -m sft-v22-q8_0.gguf --host 127.0.0.1 --port 9250 -ngl 99 -c 2048 --jinja
It is loaded only when routing is genuinely unsure, never as the default path, and released after an idle window.
On CPU alone
Measured on this architecture at this size, on an Apple M-series machine with -ngl 0 -t 4: no GPU offload at
all, four threads, and nothing in the server log naming Metal or a GPU.
| latency | tokens/s | |
|---|---|---|
| strip a negation, 8 tokens out | 38 ms | 210 |
| split one message into two requests, 11 tokens | 32 ms | 344 |
answer NO_REQUEST, 4 tokens |
17 ms | 236 |
762 MB resident while serving.
llama-server -m sft-v22-q8_0.gguf --host 127.0.0.1 --port 9250 -ngl 0 -t 4 -c 2048
What this does and does not establish: it runs usefully with no accelerator, on few threads. It has not been run on a phone or a tablet, and the three things that would decide it there are the three not measured here — 762 MB resident is a lot for a mobile process, loading 300 MB from slow storage is a different question from generating in 38 ms, and mobile cores throttle under heat while these did not.
The data it was trained on
Gramscii-IT/semantic-repair-routing
— published, and everything below can be recounted from it.
The base is that dataset's 84,819 pairs (82,819 train, 1,000 validation, 1,000 test), cleaned by deterministic rules before training: template leaks, changed values, English targets for other languages, repeated lines and generic "category" lines removed, which leaves 80,297 train, 966 validation and 963 test rows. To them this release adds 2,515 rows in Italian and English, generated by Qwen3.6-35B-A3B and kept only when every line's top capability on the engine's own roster was the one the row was written for, every number and name of the message survived, and no row shared most of its trigrams with a probe, a real question or another row. They cover a request about the conversation itself, a bare statistical topic, "what's new in X", names in lowercase or with a typo, long noisy messages carrying two or three requests, and acknowledgements that ask for nothing. Each is repeated twice in training, about 6% of it.
The base dataset, as published: Each pair is a message somebody could plausibly write and the requests inside it restated plainly, one per line.
Five languages, close to balanced:
| rows | ||
|---|---|---|
| English | 18,006 | 21.7% |
| Spanish | 16,903 | 20.4% |
| German | 16,659 | 20.1% |
| French | 16,602 | 20.0% |
| Italian | 14,649 | 17.7% |
What the message is doing, which is what the data is really organised by. Every one of the 82,819 training rows carries one of 38 such labels; these are the ten largest:
| rows | ||
|---|---|---|
massive_real — real assistant traffic |
5,736 | 6.9% |
minimal_pair — two messages differing in one word |
5,299 | 6.4% |
indirect — the request is implied, never stated |
4,912 | 5.9% |
presto_native — natively written, not translated |
4,606 | 5.6% |
oasst_human — human-written |
4,450 | 5.4% |
multi — more than one request in one message |
3,974 | 4.8% |
direct — the plain case |
3,904 | 4.7% |
vent_statement_request — a complaint with a request inside |
3,516 | 4.2% |
negation_exclusion — names what is not wanted |
3,398 | 4.1% |
trap_topic_lure — a topic named to pull routing the wrong way |
3,324 | 4.0% |
The remaining 28 cover self-correction, code-switching, ultra-short messages, typos and slang, messages carrying two, three or four requests, eight kinds of date reference — a weekday in the past, a named holiday, a bounded range — and the two families below.
The two families that decide what this model refuses to do
injection_meta — 2,278 rows — is a message aimed at the assistant
itself: reveal your instructions, ignore what you were told, pretend you
have no rules. 1,901 of them answer NO_REQUEST, because there is
nothing here to route. Together with injection_command these two families
supply 19% of every NO_REQUEST row in the set: refusing is largely
taught by messages that try hardest to get an answer.
injection_command — 2,410 rows — is the opposite lesson, and the more
delicate one. A hostile but explicit command about the user's own things:
delete every saved bank account, empty my personal cloud, cancel all my
bookings tonight. 2,281 of them are restated faithfully, action intact.
The rewriter does not sanitise and does not soften. Deciding whether an action is allowed is not its job and it has no way to do that job well: it sees one sentence, not the account, the permissions or the consequences. A model that quietly dropped "delete" would hand the router a different request from the one the person made, and the person would never learn their instruction had been edited. So the restatement carries the command as written, the router places it, and whatever runs it is where refusing belongs.
12.6% of rows (10,432) answer NO_REQUEST, and 16.9% (14,028)
answer with more than one line.
Who wrote it, and what was thrown away:
| generated by Qwen3.6-35B | 72,697 rows, 87.8% |
| generated by Qwen3.8-27B | 7,935 rows, 9.6% |
| generated by Gemma-4-E4B | 2,187 rows, 2.6% |
| rejected by the judge, not in the set | 72,996 rows |
A separate model read every message against written rules and answered a two-word verdict, decoding greedily so a verdict does not change between runs. The dataset card says where judge and generator were different models and where they stopped being.
Training
Fine-tuned from google/gemma-3-270m with mlx_lm lora, full fine-tune, all
18 blocks. The released weights are checkpoint 21,000, chosen among the
run's checkpoints by exact answers on the validation split (301 of 1,077),
never on the test split or the benches above.
Where all of this happened
One desktop machine, and nothing left it.
| machine | Mac mini, Mac16,11 |
| chip | Apple M4 Pro |
| CPU | 14 cores, 10 performance and 4 efficiency |
| GPU | 20 cores |
| memory | 64 GB unified |
| system | macOS 26.4 |
| the fine-tune | |
|---|---|
| method | mlx_lm lora, full fine-tune |
| iterations | 21,276 |
| tokens seen | 2,327,566 |
| peak memory | 5.4 GB |
| speed | about 6 iterations/s |
| wall clock | 08:27 to 09:30, about an hour |
Generators and judge ran on that same machine as local llama.cpp servers on loopback. No message here was written by a hosted API, and none was sent to one to be judged.
A model that decides where a question goes does not need a cluster or a month. It needs a narrow job, data built for it, and a bench honest enough to say when a round did not help.
Licence
This is a fine-tune of Gemma-3-270M — a Model Derivative under the Gemma Terms of Use, which it inherits whole. What those terms actually oblige, by their own sections:
Use (§3.2): the restricted uses in the Gemma Prohibited Use Policy are incorporated into the agreement by reference, and they bind anyone using these weights.
Redistribution (§3.1): whoever passes these weights on must include the §3.2 use restrictions as an enforceable provision in their own agreement, give every recipient a copy of the Terms, mark modified files as modified, and ship a notice stating "Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms". Those duties travel with the file, however many hands it passes through.
Outputs (§3.3): Google claims no rights in what the model writes. The restatements are yours, and so is the responsibility for them.
An engine that downloads these weights onto the operator's own machine, pinned by sha256, redistributes nothing: the duties above fall on whoever ships the file, not on whoever runs it.
- Downloads last month
- 599
Quantized
Model tree for Gramscii-IT/SemanticRepair-270M
Base model
google/gemma-3-270mDataset used to train Gramscii-IT/SemanticRepair-270M
Evaluation results
- Request-presence accuracy (%) on Fresh bilingual adversarial challenge (160 items)test set Gramscii author evaluation; full native traces87.500
- Request recall (%) on Fresh bilingual adversarial challenge (160 items)test set Gramscii author evaluation; full native traces87.333
- Nonrequest detection (%) on Fresh bilingual adversarial challenge (160 items)test set Gramscii author evaluation; full native traces90.000
- All annotated strings retained (%) on Fresh bilingual adversarial challenge (160 items)test set Gramscii author evaluation; full native traces81.944
- Native system exact set or correct refusal (%) on Fresh bilingual adversarial challenge (160 items)test set Gramscii author evaluation; full native traces33.750
- Native system exact multi-capability sets (%) on Fresh bilingual adversarial challenge (160 items)test set Gramscii author evaluation; full native traces2.500
- Native system unwanted capability admissions (%) — lower is better on Fresh bilingual adversarial challenge (160 items)test set Gramscii author evaluation; full native traces33.750