Viscer-4B v0.1 — research preview

This is a research preview, not a production dependency. It is a small fine-tune of Qwen3.5-4B that compresses coding-agent session transcripts into a structured JSON state. On 16 held-out sessions it reaches a median 0.712 critical-literal recall (vs 0.337 for the few-shot base model) and emits a schema-valid state in 12 of 16 cases — retry on parse failure if you need guaranteed validity. Read Known limitations before relying on it.

All measurements come from the project's frozen evaluation harness and are reproducible from the bundled technical report and tooling. Version 0.1 freezes the artifact; later versions will supersede it rather than silently change it.

What this is

Viscer is a context compressor for coding agents. It reads a long coding-agent session transcript and emits a structured JSON state object an agent can resume from: objective, hard constraints, current state, decisions (with supersessions marked), artifacts, exact tool facts, error signatures, open and completed work, rejected approaches, verification commands, and a list of verbatim anchors. Fine-tuned from Qwen3.5-4B on ledger-verified compression targets.

Files

File Size Notes
viscer-r2-step60-UD-Q4_K_XL.gguf 2.68 GiB dynamic quant (recipe below); MTP head pinned at Q8_0
viscer-r2-step60-Q4_K_M.gguf 2.78 GiB uniform Q4_K_M baseline quant
viscer-r2-step60-f16.gguf 8.07 GiB reference precision (also the merge source)
viscer-r2-step60.imatrix.gguf 3.6 MB importance matrix used for the dynamic quant
viscer-r2-tensor-types.txt — per-tensor recipe (reproducibility)

Required deployment settings (measured, not optional)

./llama-cli -m viscer-r2-step60-UD-Q4_K_XL.gguf \
  -f your_session_prompt.txt \
  -n 8192 -c 32768 -ngl 99 --temp 0 \
  --repeat-penalty 1.1 --repeat-last-n 256 \
  --jinja -st --simple-io \
  --chat-template-kwargs '{"enable_thinking": false}'
  • --repeat-penalty 1.1 --repeat-last-n 256 is required. Without it the model can enter a repetition attractor and emit tens of thousands of tokens that never close the JSON object. This is a decoding-path property of the task, not a preference the model lacks (measured margin: the model already prefers correct outputs by ~32 log-odds).
  • Size -c to the session, not to the model maximum. A 262,144-token context allocation more than halves prefill throughput.
  • Give the output room to finish: the model compresses ~5×, so a 12K-token session needs roughly 4–6K output tokens. A small cap truncates valid work.
  • Prompt = the system rules (prompts/compressor-system-v2.md) + the session transcript wrapped in the documented markers; few-shot exemplars help.

Measured quality (16 held-out cases, 14 domains, never trained on)

Metric Viscer R2 Base Qwen3.5-4B (few-shot)
parseable output 15/16 (94%) 14/16 (88%)
core schema valid 12/16 13/16
v2 schema valid (incl. handling) 12/16 1/16
critical-literal recall — median 0.712 0.337
critical-literal recall — mean 0.656 0.311
verbatim anchors — median 37 14
delivered tokens — median 2,924 1,172

Long context (R2, Q4_K_M): 144K-token session → valid output, recall 0.514, 22.8× compression, 29.4 s TTFT, 12.5 GiB peak. 167K-token session → valid, recall 0.280, 30.6× compression. No OOM at 167K on a 32 GiB card.

Speed (RTX 5090, single stream): 179 tok/s decode without speculative decoding; 212 tok/s with MTP draft depth 2 (acceptance 0.658). Depth 3 is counterproductive on this model (199 tok/s, acceptance 0.482) — the draft head was not fine-tuned, so its acceptance is lower than on the base model (0.865). ~90 s end-to-end to compress a 167K-token session.

Intended use

Compressing coding-agent session history before the next agent turn, locally, on consumer hardware. Not intended as a general summariser, a chat model, or for non-coding text.

Limitations (measured, not hypothetical)

  1. Schema validity is 12/16, not 16/16. The residual failures are JSON syntax slips (a missing colon inside otherwise complete output), not semantic errors. Callers that require guaranteed validity should retry on parse failure or use schema-constrained decoding (which guarantees structure but measured ~0.10 lower recall in our tests).
  2. Recall is ~0.71 median, so roughly one in four designated literals is lost on a typical held-out session. Treat the output as a high-quality index, not a lossless record, and keep the full transcript until the agent has confirmed the state.
  3. Trained on synthetic sessions. The corpus is synthetically generated (DeepSeek-V4-Flash) with programmatically verified fact ledgers; there has been no independent human audit. Sessions from unusual domains may degrade.
  4. Budget adherence needs a matched cap. The model writes richer states than the base model and will overshoot a tight cap rather than truncating content.
  5. Compression ceiling. Literal-complete compressions land at ~4–9× (median 5.2×). Requests for 16× compression at full literal fidelity are not achievable by this model.

Training summary

QLoRA (r=16, α=32; 21.2M trainable = 0.81%), loss masked to the target JSON only, 239 examples, 2 epochs, lr 1.5e-4, 100 minutes on one RTX 5090, peak 15.8 GiB. Data: 294 synthetic sessions (14 domains × 10 shapes, 140/140 cells), 262 verified positive targets, 88 negative/fallback targets (schema v2 handling field: compressed | passthrough | refused). Full method and 16 findings: docs/phase2-methods-and-findings.md.

License

Base model is Apache-2.0 (Qwen3.5-4B); this derivative follows the same terms. To confirm before publication.

Quantization parity (measured, same 16 held-out cases)

Metric UD-Q4_K_XL (dynamic, ours) uniform Q4_K_M
parseable 16/16 15/16
v2 / core schema 14/16 12/16
recall — median / mean 0.681 / 0.717 0.712 / 0.656
anchors — median 40 37
decode 181 tok/s 184 tok/s
peak VRAM 6,940 MiB 6,826 MiB

The dynamic recipe keeps the MTP/NextN head at Q8_0 (uniform Q4_K_M leaves it at Q4_K), which is the configuration Phase 1 recommended for speculative decoding. It is the headline release artifact.

MTP / speculative decoding — measured, and a documented dead end

The draft head was not fine-tuned, and its acceptance reflects that: 0.658 versus 0.865 on the base model, i.e. +18 % decode speed at depth 2 instead of the +42 % the base model achieves. Depth 3 is worse again (0.482), so depth 2 is the recommended setting.

Retraining the drafter in isolation was attempted twice and degraded acceptance in both cases (0.141 and 0.061). Transformers ignores the mtp.* tensors on load, so the head has to be assembled and trained by hand, and doing so optimises a hypothesis about how the runtime consumes it — a hypothesis our measurements falsified. Details in the technical report (finding F18). Anyone attempting this should read llama.cpp's draft-mtp graph construction first and treat the runtime as the specification.

v0.1 release contents

Item Status
viscer-r2-step60-UD-Q4_K_XL.gguf (sha256 44e0ae622a9172c1…) ready
viscer-r2-step60-Q4_K_M.gguf, -f16.gguf, imatrix, tensor recipe ready
This model card ready
Technical report (docs/phase2-methods-and-findings.md, 18 findings) ready
Decisions log (DECISIONS.md, D1–D46) ready
Synthetic corpus + ledgers (294 sessions) decision pending: publish for reproducibility, or hold
Independent human audit of ledgers not done — disclosed as a limitation
Phase-1 benchmark report (base-model selection) ready

Known limitations that matter for use

  1. Schema validity is 12/16 on held-out cases against a 99.5 % release target; the residual failures are JSON syntax slips, not semantic errors. Retry on parse failure.
  2. Recall ~0.71 median — a high-quality index, not a lossless record. Keep the transcript until the agent confirms the state.
  3. Synthetic training corpus, programmatically verified, no independent human audit.
  4. The compression ceiling is ~4–9×; 16× at full literal fidelity is not achievable.
  • MTP acceptance measurement on the fine-tuned model
  • Independent human spot-audit of the synthetic ledgers
  • Hugging Face repository, tags, and paper link
Downloads last month
232
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TheWickedGustav/viscer-4b

Finetuned
Qwen/Qwen3.5-4B
Quantized
(481)
this model