Instructions to use AtomicChat/d1-3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AtomicChat/d1-3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/d1-3B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/d1-3B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/d1-3B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/d1-3B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AtomicChat/d1-3B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf AtomicChat/d1-3B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AtomicChat/d1-3B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf AtomicChat/d1-3B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/AtomicChat/d1-3B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use AtomicChat/d1-3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AtomicChat/d1-3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/d1-3B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/AtomicChat/d1-3B-GGUF:Q4_K_M
- Ollama
How to use AtomicChat/d1-3B-GGUF with Ollama:
ollama run hf.co/AtomicChat/d1-3B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use AtomicChat/d1-3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AtomicChat/d1-3B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AtomicChat/d1-3B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use AtomicChat/d1-3B-GGUF with Docker Model Runner:
docker model run hf.co/AtomicChat/d1-3B-GGUF:Q4_K_M
- Lemonade
How to use AtomicChat/d1-3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AtomicChat/d1-3B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.d1-3B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use AtomicChat/d1-3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AtomicChat/d1-3B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AtomicChat/d1-3B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AtomicChat/d1-3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AtomicChat/d1-3B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AtomicChat/d1-3B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
How to Run d1-3B Locally
Built from Liquid AI's published weights with our own importance matrix, and measured on decisions, not just text.
d1-3B is Liquid AI's decision model. It takes a state and answers typed questions in one forward pass: yes or no, a pick from named options, or a score. The answer is read from the model's distribution over the options' tokens, so nothing is generated.
Pick a file
Every number below is measured on one machine. The raw results and logs are in the metrics repo.
same answer: across 1,285 decisions, how often the file gives the same answer as the BF16 weights published in LiquidAI/d1-3B. The questions are yes/no, choice and score questions over held-out texts in 30 languages and source code. The number of changed answers is in brackets.option drift: the mean total variation distance between the file's option probabilities and the original's. 0 means identical.KLDandtop-1: full-vocabulary agreement with the original on held-out text.
| File | Size | same answer | option drift | KLD | top-1 |
|---|---|---|---|---|---|
BF16 |
5,403 MB on disk | reference | 0 | 0 | 100% |
Q8_0 |
2,875 MB on disk | 99.3% (9) | 0.0033 | 0.0010 | 98.28% |
AD-Q6_K |
2,348 MB on disk | 99.1% (11) | 0.0053 | 0.0022 | 97.41% |
AD-Q5_K_M |
1,953 MB on disk | 98.1% (25) | 0.0112 | 0.0080 | 95.09% |
AD-Q4_K_M |
1,658 MB on disk | 97.1% (37) | 0.0178 | 0.0244 | 91.54% |
AD-IQ4_XS |
1,570 MB on disk | 95.6% (57) | 0.0203 | 0.0283 | 90.92% |
AD- means Atomic Dynamic: the type is chosen per tensor instead of taken from a llama.cpp preset.
- The token table, which d1-3B also uses as its output head, never goes below Q6_K.
- Attention stays at Q8_0 or Q6_K.
ffn_downin the first and last four blocks takes one step more than the middle.
The projector for images comes as mmproj-d1-3B-BF16 and mmproj-d1-3B-Q8_0. We did not measure
image decisions.
How these compare to other GGUFs
- Liquid's own GGUFs are built from the current weights and
carry the type
lfm2-d1. Their BF16 has the same tensors as ours. - prithivMLmods/d1-3B-GGUF has llama.cpp's stock presets, also from the current weights.
We measured all of them on one machine (an NVIDIA A10, llama.cpp 88dcc46) against our BF16. This run is
separate from the table above, so our own rows differ from it slightly.
| File | Size | same answer | option drift | option KL | KLD | top-1 |
|---|---|---|---|---|---|---|
AtomicChat Q8_0 |
2,875 MB | 99.2% (10) | 0.0034 | 0.0001 | 0.0009 | 98.30% |
Liquid Q8_0 |
2,875 MB | 99.4% (8) | 0.0034 | 0.0001 | 0.0009 | 98.30% |
AtomicChat AD-Q6_K |
2,348 MB | 99.3% (9) | 0.0054 | 0.0002 | 0.0022 | 97.41% |
prithivMLmods Q6_K |
2,222 MB | 98.7% (17) | 0.0077 | 0.0004 | 0.0041 | 96.34% |
AtomicChat AD-Q5_K_M |
1,953 MB | 97.7% (29) | 0.0112 | 0.0009 | 0.0081 | 95.11% |
prithivMLmods Q5_K_M |
1,940 MB | 97.7% (29) | 0.0138 | 0.0014 | 0.0121 | 93.89% |
AtomicChat AD-Q4_K_M |
1,658 MB | 97.0% (38) | 0.0174 | 0.0023 | 0.0243 | 91.55% |
Liquid Q4_K_M |
1,674 MB | 95.7% (55) | 0.0270 | 0.0053 | 0.0392 | 89.42% |
prithivMLmods Q4_K_M |
1,674 MB | 95.9% (53) | 0.0272 | 0.0053 | 0.0389 | 89.44% |
- 4 bits. Our
AD-Q4_K_Mis 16 MB smaller than the two stockQ4_K_Mfiles.- It changes 38 answers instead of 55.
- Its option KL is less than half of theirs, and its KLD is 38% lower.
- 5 bits. At about the same size as prithivMLmods'
Q5_K_M, ours changes the same number of answers, with about a third less option KL and KLD. - 6 bits. Our
AD-Q6_Kis 126 MB larger than prithivMLmods'Q6_K, so the two are not a same-size comparison. - Q8_0. The two Q8_0 files measure the same.
- A few tensors round differently, because Liquid's converter writes Q8_0 itself.
- 8 against 10 changed answers is within run-to-run noise. Our own
Q8_0changed 9 in the run above.
Liquid's first GGUFs, published on 6 October, were built from the previous version of the weights. Liquid replaced them on 7 October. Our measurement of those first files is in the metrics repo.
Running it
This needs a llama.cpp build that includes #30110: commit 88dcc46, merged on 7 October, or newer. Older
builds do not know the type lfm2-d1 and refuse to load the file.
llama-server -m d1-3B-AD-Q4_K_M.gguf --mmproj mmproj-d1-3B-Q8_0.gguf -ngl 99 -c 8192
Ask through the /v1/systemone endpoint. It takes a state and named, typed questions, and returns each answer
with its option probabilities:
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
"state": "I was charged twice this month, please refund one of them.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'
Response from AD-Q4_K_M, numbers rounded:
{
"answers": {
"team": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.984, "technical": 0.011, "fraud": 0.005}, "confidence": 0.977},
"angry": {"type": "noul", "noul": 0.099}
},
"usage": {"input_tokens": 97, "output_tokens": 0}
}
noulquestions are yes or no, andscorequestions take a list of 2 to 10 levels.- Images go in an
imagesarray as data URLs. - The full request format is in the server documentation, and the question schema in the original card.
Checked on the merged endpoint. We ran the PR's final commit against Liquid's own PyTorch code in FP32, on the same 1,285 text decisions:
Q8_0changes 6 answers, with an option KL of 0.0001.AD-Q4_K_Mchanges 44, with an option KL of 0.0022.- These numbers use a different reference and a different readout from the table above, so they do not match it exactly.
How we made and measured them
- BF16: converted from LiquidAI/d1-3B (revision
da1fe36) with llama.cpp18b5f8b. It matches Liquid's PyTorch code on 300 decisions with a mean option KL of 0.0002. The 9 changed answers there are near-ties, with a largest option difference of 0.035. The projector is tensor-for-tensor identical to Liquid's. - Importance matrix: 1.5M tokens of decision prompts, 3,206 in all. The states come from the calibration corpora pool (Wikipedia in 30 languages, code, structured files). Each carries yes/no, choice and score questions and is rendered with the model's own prompt code.
- Decisions: 1,285 questions over the held-out calib-corpora
eval/neutralandeval/codetexts, none of them in the calibration set. Before the endpoint existed, they were answered from the option tokens' log-probabilities on/completion: the prompt is rendered with Liquid'sprompt.py, and the readout is a softmax over the options' tokens. The scripts are in the metrics repo. - Text:
llama-perplexity --kl-divergenceovereval/neutral, 93 chunks of 4,096 tokens. - Setup: an NVIDIA H100 and llama.cpp
18b5f8bfor every file. - Metadata: on 7 October, after #30110 was merged, the files were re-stamped. The text files got
lfm2.decision.type = lfm2-d1and thesystemonetemplate, both taken from a BF16 file the PR's converter wrote. The projectors gotclip.vision.image_resize_algo = bicubic. No tensor changed.
Model details
- Base: LiquidAI/d1-3B, a decision model post-trained from LFM2.5-VL-3B: 3.1B parameters, 32K context, a 128K vocabulary and a SigLIP2 vision encoder.
- License: LFM Open License v1.0; see
LICENSE.
- Downloads last month
- -
4-bit
5-bit
6-bit
8-bit
16-bit


