Consylr 4B

An offline exception analyst for enterprise record reconciliation. Given a packet of records, policy lines and documents, it returns one JSON object: which records correspond, which differences are exceptions and why, what remains unmatched, and what a human should do next.

It declines. Where a difference is real but the packet does not explain it, the model says cause_not_established, cites nothing, and opens a request for the evidence it would need. That behaviour is the point of the model, and it is measured rather than asserted — see Evaluation below.

Specification

Parameters 4.02 B
Quantisation Q6_K, importance-matrix calibrated on the training distribution
File size 3,306,257,184 bytes (3.08 GiB)
Context used 2,048 tokens (the file declares 262,144 — see below)
Runtime llama.cpp b10453, CPU-only, -ngl 0
Interface raw completion only; not conversational
Language English only

Files

file what it is
consylr-v5-Q6_K-imat.gguf the weights, sha256 43c5205d773e7262…
consylr-v5-Q6_K-imat.manifest.json provenance chain, every digest computed rather than transcribed
README.md this card

Run it

Raw completion only. This model is not conversational. Hugging Face may show a chat-template badge because the file carries a template; do not use it. See Chat mode below.

llama-completion -m consylr-v5-Q6_K-imat.gguf \
  -c 2048 -n 320 --temp 0 -ngl 0 -no-cnv --no-display-prompt \
  -f your-packet-prompt.txt

-c 2048 is deliberate: the file declares a 262,144-token context and -c defaults to the model's own, which asks for a KV cache far past any laptop.

Prompt format

No chat templating is applied — raw completion. The file does carry a chat template; it is unused and must not be applied (see "Chat mode aborts" below). There is no system turn and no role markers. The model was trained on one frozen template with the packet substituted into it, and it must be given exactly that string — nothing before it, nothing between the prompt and the answer. The shape is:

<instructions — the frozen preamble, byte-identical to training>

RECORDS
<the ledger rows under comparison>

POLICY
<the policy lines that may or may not explain a difference>

DOCUMENTS
<supporting documents, where present>

ANSWER

The model continues from ANSWER and emits a single JSON object. Two properties follow from this and are worth knowing before you wire anything up:

  • An over-budget packet must be refused, not truncated. The harness reserves the generation budget and left-truncates the prompt, which yields well-formed JSON about records the model was never shown. That failure is invisible downstream.
  • Wrapping the prompt in a chat template changes the string and therefore the behaviour. See below — it does not silently degrade, it aborts.

Chat mode aborts, and that is intended

Conversation mode auto-enables whenever a GGUF carries a chat template. This one does, and applying it raises an uncaught exception in common_chat_templates_apply — the process aborts at load, before generating a single token.

This is a property worth having rather than a defect to work around: there is no path where a chat wrapper produces plausible-looking wrong answers, because it never produces any. Pass -no-cnv and the correct interface is the only one that runs.

Evaluation

Measured on a held-out sealed set of 204 packets across fourteen families, and on a twin probe of 75 minimal pairs — two packets identical in every record, amount, date and document, differing by one policy line.

Exact match 45.6%
Correspondence F1 0.881 (positional null control: 0.839)
Schema valid 90.2%
Citation resolution 100%
Correct abstention 94.7%
Abstention-direction flip rate 77.3% (58 of 75)
Strict flip rate (schema-valid, correct record, correct correspondences) 49.3% (37 of 75)
Declined when it should not have 0 of 75 on the probe; 27.2% on the sealed set

Every F1 is reported beside a null control — a model that reads nothing and pairs record n with record n scores 0.839 on this exam, so the raw figure alone would flatter any system.

Abstention has two directions and both are printed. The two figures in the last row are different instruments, not a contradiction. The probe's settled halves are twins of withheld ones, so evidence-presence is the only variable that moves; there the model never declines wrongly. The sealed set is the harder mixed distribution, and there it over-declines on 27.2% of establishable cases — well above the 5% we were aiming at. The gap between 0 and 27.2% is the finding: when evidence-presence is the only thing that changes, the model reads it correctly; under distribution pressure it errs toward declining. That is the direction of error we chose, because the alternative is a confident fabricated cause.

The abstention-direction flip rate measures the direction of the decision, not the full correctness of either answer. A pair counts when the settled packet is answered and the withheld one declines for an unestablished cause. Aggregate abstention scores can be earned by a model that declines everything or asserts everything; a model that changes its answer when one policy line is removed from an otherwise identical packet is reading evidence, and nothing else explains it.

It does not check which record the declined cause was attached to, whether the output was schema-valid, or whether the correspondences were right. Scored strictly on all three, the same runs give 37 of 75 rather than 58: of the 58 that pass on direction, 11 name the wrong record and 7 have a schema-invalid half, and only 3 pairs are exact on both complete outputs. The direction rate is what separates reading a policy line from reading the family; the strict rate is what to assume about end-to-end correctness.

That distinction is not academic. An earlier quantisation of these same weights scored 1.3% on the probe while the sealed set still reported 75.4% correct abstention — the exam could not see the loss, because it contains no withheld twins. The importance matrix used here was calibrated on the training distribution specifically to preserve that margin.

Against the base model

Same sealed set, same 204 packets.

base checkpoint this model
Parse rate 89.7% 100%
Schema valid 43.1% 90.2%
Correspondence F1 0.433 0.881
Exact match 0.0% 45.6%
Abstention declines every case (correct abstention 100%, false abstention 100%) 94.7% correct, 27.2% false

The base column uses a scaffolded harness. Under the deployment interface itself — bare completion, no scaffold — the base checkpoint produces no parseable output at all, which is why the comparison scaffolds it rather than reporting zeros.

Read the abstention cell carefully: 100% correct abstention is not a good score, it is what declining everything looks like, and it is exactly why that figure is never printed without its false-abstention twin. The base checkpoint can be coaxed into producing structure. It cannot be coaxed into judgement — that is what the fine-tuning bought.

Provenance

consylr-v5-Q6_K-imat.manifest.json ships beside the weights and records every digest — computed, never transcribed — back through the unquantised parent and the adapter to the training corpus.

Artefact consylr-v5-Q6_K-imat.gguf, sha256 43c5205d773e7262…
Quantisation Q6_K with an importance matrix calibrated on the training distribution
Runtime llama.cpp b10453
Base mlx-community/Qwen3-4B-Instruct-2507-4bit at 50d427756c…, Apache-2.0

Limits

  • English only, and enterprise reconciliation only. It is not a general assistant and has no conversational ability.
  • Training data is synthetic, generated procedurally. It has never seen a real ledger.
  • Arithmetic is not the model's job. It characterises a difference; the amounts are verified deterministically outside it.
  • It over-declines under distribution pressure — 27.2% false abstention on the sealed set. Budget for a human reviewing declined cases.
  • The abstention behaviour is measured on twins built by the same generator as the training data. That measures generalisation to unseen instances of a construction the model has seen — not a general capability to know what it does not know.

Licence

Apache-2.0, inherited from the base checkpoint mlx-community/Qwen3-4B-Instruct-2507-4bit (verified at revision 50d427756c6b1b2fe0c0a10f67fbda1fc8e82c1b, whose own base Qwen/Qwen3-4B-Instruct-2507 is likewise Apache-2.0).

Downloads last month
11
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for badrama/consylr-4b

Quantized
(4)
this model

Evaluation results

  • Exact match on sealed-v4 — 204 held-out packets, 14 families
    self-reported
    0.456
  • Correspondence F1 (positional null control 0.839) on sealed-v4 — 204 held-out packets, 14 families
    self-reported
    0.881
  • Schema valid on sealed-v4 — 204 held-out packets, 14 families
    self-reported
    0.902
  • Citation resolution on sealed-v4 — 204 held-out packets, 14 families
    self-reported
    1.000
  • Correct abstention on sealed-v4 — 204 held-out packets, 14 families
    self-reported
    0.947
  • False abstention (lower is better) on sealed-v4 — 204 held-out packets, 14 families
    self-reported
    0.272
  • Abstention-direction flip rate on probe-v1 — 75 minimal pairs
    self-reported
    0.773