Instructions to use TheWickedGustav/viscer-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use TheWickedGustav/viscer-4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL
Use Docker
docker model run hf.co/TheWickedGustav/viscer-4b:UD-Q4_K_XL
- LM Studio
- Jan
- vLLM
How to use TheWickedGustav/viscer-4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TheWickedGustav/viscer-4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheWickedGustav/viscer-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/TheWickedGustav/viscer-4b:UD-Q4_K_XL
- Ollama
How to use TheWickedGustav/viscer-4b with Ollama:
ollama run hf.co/TheWickedGustav/viscer-4b:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use TheWickedGustav/viscer-4b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TheWickedGustav/viscer-4b:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use TheWickedGustav/viscer-4b with Docker Model Runner:
docker model run hf.co/TheWickedGustav/viscer-4b:UD-Q4_K_XL
- Lemonade
How to use TheWickedGustav/viscer-4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull TheWickedGustav/viscer-4b:UD-Q4_K_XL
Run and chat with the model
lemonade run user.viscer-4b-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use TheWickedGustav/viscer-4b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TheWickedGustav/viscer-4b:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TheWickedGustav/viscer-4b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TheWickedGustav/viscer-4b:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TheWickedGustav/viscer-4b:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Viscer-4B v0.1 — research preview
- What this is
- Files
- Required deployment settings (measured, not optional)
- Measured quality (16 held-out cases, 14 domains, never trained on)
- Intended use
- Limitations (measured, not hypothetical)
- Training summary
- License
- Quantization parity (measured, same 16 held-out cases)
- MTP / speculative decoding — measured, and a documented dead end
- v0.1 release contents
- Known limitations that matter for use
- What this is
Viscer-4B v0.1 — research preview
This is a research preview, not a production dependency. It is a small fine-tune of Qwen3.5-4B that compresses coding-agent session transcripts into a structured JSON state. On 16 held-out sessions it reaches a median 0.712 critical-literal recall (vs 0.337 for the few-shot base model) and emits a schema-valid state in 12 of 16 cases — retry on parse failure if you need guaranteed validity. Read Known limitations before relying on it.
All measurements come from the project's frozen evaluation harness and are reproducible from the bundled technical report and tooling. Version 0.1 freezes the artifact; later versions will supersede it rather than silently change it.
What this is
Viscer is a context compressor for coding agents. It reads a long coding-agent session transcript and emits a structured JSON state object an agent can resume from: objective, hard constraints, current state, decisions (with supersessions marked), artifacts, exact tool facts, error signatures, open and completed work, rejected approaches, verification commands, and a list of verbatim anchors. Fine-tuned from Qwen3.5-4B on ledger-verified compression targets.
Files
| File | Size | Notes |
|---|---|---|
viscer-r2-step60-UD-Q4_K_XL.gguf |
2.68 GiB | dynamic quant (recipe below); MTP head pinned at Q8_0 |
viscer-r2-step60-Q4_K_M.gguf |
2.78 GiB | uniform Q4_K_M baseline quant |
viscer-r2-step60-f16.gguf |
8.07 GiB | reference precision (also the merge source) |
viscer-r2-step60.imatrix.gguf |
3.6 MB | importance matrix used for the dynamic quant |
viscer-r2-tensor-types.txt |
— | per-tensor recipe (reproducibility) |
Required deployment settings (measured, not optional)
./llama-cli -m viscer-r2-step60-UD-Q4_K_XL.gguf \
-f your_session_prompt.txt \
-n 8192 -c 32768 -ngl 99 --temp 0 \
--repeat-penalty 1.1 --repeat-last-n 256 \
--jinja -st --simple-io \
--chat-template-kwargs '{"enable_thinking": false}'
--repeat-penalty 1.1 --repeat-last-n 256is required. Without it the model can enter a repetition attractor and emit tens of thousands of tokens that never close the JSON object. This is a decoding-path property of the task, not a preference the model lacks (measured margin: the model already prefers correct outputs by ~32 log-odds).- Size
-cto the session, not to the model maximum. A 262,144-token context allocation more than halves prefill throughput. - Give the output room to finish: the model compresses ~5×, so a 12K-token session needs roughly 4–6K output tokens. A small cap truncates valid work.
- Prompt = the system rules (
prompts/compressor-system-v2.md) + the session transcript wrapped in the documented markers; few-shot exemplars help.
Measured quality (16 held-out cases, 14 domains, never trained on)
| Metric | Viscer R2 | Base Qwen3.5-4B (few-shot) |
|---|---|---|
| parseable output | 15/16 (94%) | 14/16 (88%) |
| core schema valid | 12/16 | 13/16 |
v2 schema valid (incl. handling) |
12/16 | 1/16 |
| critical-literal recall — median | 0.712 | 0.337 |
| critical-literal recall — mean | 0.656 | 0.311 |
| verbatim anchors — median | 37 | 14 |
| delivered tokens — median | 2,924 | 1,172 |
Long context (R2, Q4_K_M): 144K-token session → valid output, recall 0.514, 22.8× compression, 29.4 s TTFT, 12.5 GiB peak. 167K-token session → valid, recall 0.280, 30.6× compression. No OOM at 167K on a 32 GiB card.
Speed (RTX 5090, single stream): 179 tok/s decode without speculative decoding; 212 tok/s with MTP draft depth 2 (acceptance 0.658). Depth 3 is counterproductive on this model (199 tok/s, acceptance 0.482) — the draft head was not fine-tuned, so its acceptance is lower than on the base model (0.865). ~90 s end-to-end to compress a 167K-token session.
Intended use
Compressing coding-agent session history before the next agent turn, locally, on consumer hardware. Not intended as a general summariser, a chat model, or for non-coding text.
Limitations (measured, not hypothetical)
- Schema validity is 12/16, not 16/16. The residual failures are JSON syntax slips (a missing colon inside otherwise complete output), not semantic errors. Callers that require guaranteed validity should retry on parse failure or use schema-constrained decoding (which guarantees structure but measured ~0.10 lower recall in our tests).
- Recall is ~0.71 median, so roughly one in four designated literals is lost on a typical held-out session. Treat the output as a high-quality index, not a lossless record, and keep the full transcript until the agent has confirmed the state.
- Trained on synthetic sessions. The corpus is synthetically generated (DeepSeek-V4-Flash) with programmatically verified fact ledgers; there has been no independent human audit. Sessions from unusual domains may degrade.
- Budget adherence needs a matched cap. The model writes richer states than the base model and will overshoot a tight cap rather than truncating content.
- Compression ceiling. Literal-complete compressions land at ~4–9× (median 5.2×). Requests for 16× compression at full literal fidelity are not achievable by this model.
Training summary
QLoRA (r=16, α=32; 21.2M trainable = 0.81%), loss masked to the target JSON only,
239 examples, 2 epochs, lr 1.5e-4, 100 minutes on one RTX 5090, peak 15.8 GiB.
Data: 294 synthetic sessions (14 domains × 10 shapes, 140/140 cells), 262 verified
positive targets, 88 negative/fallback targets (schema v2 handling field:
compressed | passthrough | refused). Full method and 16 findings:
docs/phase2-methods-and-findings.md.
License
Base model is Apache-2.0 (Qwen3.5-4B); this derivative follows the same terms. To confirm before publication.
Quantization parity (measured, same 16 held-out cases)
| Metric | UD-Q4_K_XL (dynamic, ours) |
uniform Q4_K_M |
|---|---|---|
| parseable | 16/16 | 15/16 |
| v2 / core schema | 14/16 | 12/16 |
| recall — median / mean | 0.681 / 0.717 | 0.712 / 0.656 |
| anchors — median | 40 | 37 |
| decode | 181 tok/s | 184 tok/s |
| peak VRAM | 6,940 MiB | 6,826 MiB |
The dynamic recipe keeps the MTP/NextN head at Q8_0 (uniform Q4_K_M leaves it at Q4_K), which is the configuration Phase 1 recommended for speculative decoding. It is the headline release artifact.
MTP / speculative decoding — measured, and a documented dead end
The draft head was not fine-tuned, and its acceptance reflects that: 0.658 versus 0.865 on the base model, i.e. +18 % decode speed at depth 2 instead of the +42 % the base model achieves. Depth 3 is worse again (0.482), so depth 2 is the recommended setting.
Retraining the drafter in isolation was attempted twice and degraded acceptance
in both cases (0.141 and 0.061). Transformers ignores the mtp.* tensors on
load, so the head has to be assembled and trained by hand, and doing so optimises
a hypothesis about how the runtime consumes it — a hypothesis our measurements
falsified. Details in the technical report (finding F18). Anyone attempting this
should read llama.cpp's draft-mtp graph construction first and treat the runtime
as the specification.
v0.1 release contents
| Item | Status |
|---|---|
viscer-r2-step60-UD-Q4_K_XL.gguf (sha256 44e0ae622a9172c1…) |
ready |
viscer-r2-step60-Q4_K_M.gguf, -f16.gguf, imatrix, tensor recipe |
ready |
| This model card | ready |
Technical report (docs/phase2-methods-and-findings.md, 18 findings) |
ready |
Decisions log (DECISIONS.md, D1–D46) |
ready |
| Synthetic corpus + ledgers (294 sessions) | decision pending: publish for reproducibility, or hold |
| Independent human audit of ledgers | not done — disclosed as a limitation |
| Phase-1 benchmark report (base-model selection) | ready |
Known limitations that matter for use
- Schema validity is 12/16 on held-out cases against a 99.5 % release target; the residual failures are JSON syntax slips, not semantic errors. Retry on parse failure.
- Recall ~0.71 median — a high-quality index, not a lossless record. Keep the transcript until the agent confirms the state.
- Synthetic training corpus, programmatically verified, no independent human audit.
- The compression ceiling is ~4–9×; 16× at full literal fidelity is not achievable.
- MTP acceptance measurement on the fine-tuned model
- Independent human spot-audit of the synthetic ledgers
- Hugging Face repository, tags, and paper link
- Downloads last month
- 232