Instructions to use badrama/consylr-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badrama/consylr-4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badrama/consylr-4b:Q6_K # Run inference directly in the terminal: llama cli -hf badrama/consylr-4b:Q6_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badrama/consylr-4b:Q6_K # Run inference directly in the terminal: llama cli -hf badrama/consylr-4b:Q6_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badrama/consylr-4b:Q6_K # Run inference directly in the terminal: ./llama-cli -hf badrama/consylr-4b:Q6_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badrama/consylr-4b:Q6_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf badrama/consylr-4b:Q6_K
Use Docker
docker model run hf.co/badrama/consylr-4b:Q6_K
- LM Studio
- Jan
- vLLM
How to use badrama/consylr-4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badrama/consylr-4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badrama/consylr-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badrama/consylr-4b:Q6_K
- Ollama
How to use badrama/consylr-4b with Ollama:
ollama run hf.co/badrama/consylr-4b:Q6_K
- Unsloth Desktop
- Docker Model Runner
How to use badrama/consylr-4b with Docker Model Runner:
docker model run hf.co/badrama/consylr-4b:Q6_K
- Lemonade
How to use badrama/consylr-4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badrama/consylr-4b:Q6_K
Run and chat with the model
lemonade run user.consylr-4b-Q6_K
List all available models
lemonade list
- Atomic Chat
Consylr 4B
An offline exception analyst for enterprise record reconciliation. Given a packet of records, policy lines and documents, it returns one JSON object: which records correspond, which differences are exceptions and why, what remains unmatched, and what a human should do next.
It declines. Where a difference is real but the packet does not explain it,
the model says cause_not_established, cites nothing, and opens a request for
the evidence it would need. That behaviour is the point of the model, and it is
measured rather than asserted — see Evaluation below.
Specification
| Parameters | 4.02 B |
| Quantisation | Q6_K, importance-matrix calibrated on the training distribution |
| File size | 3,306,257,184 bytes (3.08 GiB) |
| Context used | 2,048 tokens (the file declares 262,144 — see below) |
| Runtime | llama.cpp b10453, CPU-only, -ngl 0 |
| Interface | raw completion only; not conversational |
| Language | English only |
Files
| file | what it is |
|---|---|
consylr-v5-Q6_K-imat.gguf |
the weights, sha256 43c5205d773e7262… |
consylr-v5-Q6_K-imat.manifest.json |
provenance chain, every digest computed rather than transcribed |
README.md |
this card |
Run it
Raw completion only. This model is not conversational. Hugging Face may show a chat-template badge because the file carries a template; do not use it. See Chat mode below.
llama-completion -m consylr-v5-Q6_K-imat.gguf \
-c 2048 -n 320 --temp 0 -ngl 0 -no-cnv --no-display-prompt \
-f your-packet-prompt.txt
-c 2048 is deliberate: the file declares a 262,144-token context and -c
defaults to the model's own, which asks for a KV cache far past any laptop.
Prompt format
No chat templating is applied — raw completion. The file does carry a chat template; it is unused and must not be applied (see "Chat mode aborts" below). There is no system turn and no role markers. The model was trained on one frozen template with the packet substituted into it, and it must be given exactly that string — nothing before it, nothing between the prompt and the answer. The shape is:
<instructions — the frozen preamble, byte-identical to training>
RECORDS
<the ledger rows under comparison>
POLICY
<the policy lines that may or may not explain a difference>
DOCUMENTS
<supporting documents, where present>
ANSWER
The model continues from ANSWER and emits a single JSON object. Two properties
follow from this and are worth knowing before you wire anything up:
- An over-budget packet must be refused, not truncated. The harness reserves the generation budget and left-truncates the prompt, which yields well-formed JSON about records the model was never shown. That failure is invisible downstream.
- Wrapping the prompt in a chat template changes the string and therefore the behaviour. See below — it does not silently degrade, it aborts.
Chat mode aborts, and that is intended
Conversation mode auto-enables whenever a GGUF carries a chat template. This one
does, and applying it raises an uncaught exception in common_chat_templates_apply
— the process aborts at load, before generating a single token.
This is a property worth having rather than a defect to work around: there is no
path where a chat wrapper produces plausible-looking wrong answers, because it
never produces any. Pass -no-cnv and the correct interface is the only one that
runs.
Evaluation
Measured on a held-out sealed set of 204 packets across fourteen families, and on a twin probe of 75 minimal pairs — two packets identical in every record, amount, date and document, differing by one policy line.
| Exact match | 45.6% |
| Correspondence F1 | 0.881 (positional null control: 0.839) |
| Schema valid | 90.2% |
| Citation resolution | 100% |
| Correct abstention | 94.7% |
| Abstention-direction flip rate | 77.3% (58 of 75) |
| Strict flip rate (schema-valid, correct record, correct correspondences) | 49.3% (37 of 75) |
| Declined when it should not have | 0 of 75 on the probe; 27.2% on the sealed set |
Every F1 is reported beside a null control — a model that reads nothing and pairs record n with record n scores 0.839 on this exam, so the raw figure alone would flatter any system.
Abstention has two directions and both are printed. The two figures in the last row are different instruments, not a contradiction. The probe's settled halves are twins of withheld ones, so evidence-presence is the only variable that moves; there the model never declines wrongly. The sealed set is the harder mixed distribution, and there it over-declines on 27.2% of establishable cases — well above the 5% we were aiming at. The gap between 0 and 27.2% is the finding: when evidence-presence is the only thing that changes, the model reads it correctly; under distribution pressure it errs toward declining. That is the direction of error we chose, because the alternative is a confident fabricated cause.
The abstention-direction flip rate measures the direction of the decision, not the full correctness of either answer. A pair counts when the settled packet is answered and the withheld one declines for an unestablished cause. Aggregate abstention scores can be earned by a model that declines everything or asserts everything; a model that changes its answer when one policy line is removed from an otherwise identical packet is reading evidence, and nothing else explains it.
It does not check which record the declined cause was attached to, whether the output was schema-valid, or whether the correspondences were right. Scored strictly on all three, the same runs give 37 of 75 rather than 58: of the 58 that pass on direction, 11 name the wrong record and 7 have a schema-invalid half, and only 3 pairs are exact on both complete outputs. The direction rate is what separates reading a policy line from reading the family; the strict rate is what to assume about end-to-end correctness.
That distinction is not academic. An earlier quantisation of these same weights scored 1.3% on the probe while the sealed set still reported 75.4% correct abstention — the exam could not see the loss, because it contains no withheld twins. The importance matrix used here was calibrated on the training distribution specifically to preserve that margin.
Against the base model
Same sealed set, same 204 packets.
| base checkpoint | this model | |
|---|---|---|
| Parse rate | 89.7% | 100% |
| Schema valid | 43.1% | 90.2% |
| Correspondence F1 | 0.433 | 0.881 |
| Exact match | 0.0% | 45.6% |
| Abstention | declines every case (correct abstention 100%, false abstention 100%) | 94.7% correct, 27.2% false |
The base column uses a scaffolded harness. Under the deployment interface itself — bare completion, no scaffold — the base checkpoint produces no parseable output at all, which is why the comparison scaffolds it rather than reporting zeros.
Read the abstention cell carefully: 100% correct abstention is not a good score, it is what declining everything looks like, and it is exactly why that figure is never printed without its false-abstention twin. The base checkpoint can be coaxed into producing structure. It cannot be coaxed into judgement — that is what the fine-tuning bought.
Provenance
consylr-v5-Q6_K-imat.manifest.json ships beside the weights and records every
digest — computed, never transcribed — back through the unquantised parent and
the adapter to the training corpus.
| Artefact | consylr-v5-Q6_K-imat.gguf, sha256 43c5205d773e7262… |
| Quantisation | Q6_K with an importance matrix calibrated on the training distribution |
| Runtime | llama.cpp b10453 |
| Base | mlx-community/Qwen3-4B-Instruct-2507-4bit at 50d427756c…, Apache-2.0 |
Limits
- English only, and enterprise reconciliation only. It is not a general assistant and has no conversational ability.
- Training data is synthetic, generated procedurally. It has never seen a real ledger.
- Arithmetic is not the model's job. It characterises a difference; the amounts are verified deterministically outside it.
- It over-declines under distribution pressure — 27.2% false abstention on the sealed set. Budget for a human reviewing declined cases.
- The abstention behaviour is measured on twins built by the same generator as the training data. That measures generalisation to unseen instances of a construction the model has seen — not a general capability to know what it does not know.
Licence
Apache-2.0, inherited from the base checkpoint
mlx-community/Qwen3-4B-Instruct-2507-4bit (verified at revision
50d427756c6b1b2fe0c0a10f67fbda1fc8e82c1b, whose own base
Qwen/Qwen3-4B-Instruct-2507 is likewise Apache-2.0).
- Downloads last month
- 11
6-bit
Model tree for badrama/consylr-4b
Base model
Qwen/Qwen3-4B-Instruct-2507Evaluation results
- Exact match on sealed-v4 — 204 held-out packets, 14 familiesself-reported0.456
- Correspondence F1 (positional null control 0.839) on sealed-v4 — 204 held-out packets, 14 familiesself-reported0.881
- Schema valid on sealed-v4 — 204 held-out packets, 14 familiesself-reported0.902
- Citation resolution on sealed-v4 — 204 held-out packets, 14 familiesself-reported1.000
- Correct abstention on sealed-v4 — 204 held-out packets, 14 familiesself-reported0.947
- False abstention (lower is better) on sealed-v4 — 204 held-out packets, 14 familiesself-reported0.272
- Abstention-direction flip rate on probe-v1 — 75 minimal pairsself-reported0.773