How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Snapkitty/snapkitty-merged:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Snapkitty/snapkitty-merged:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Snapkitty/snapkitty-merged:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Snapkitty/snapkitty-merged:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Snapkitty/snapkitty-merged:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf Snapkitty/snapkitty-merged:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Snapkitty/snapkitty-merged:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Snapkitty/snapkitty-merged:Q4_K_M
Use Docker
docker model run hf.co/Snapkitty/snapkitty-merged:Q4_K_M
Quick Links

snapkitty-merged — Nemotron 4.2B Merged Q4_K_M

Merged and fine-tuned Nemotron Mini 4.2B, quantized to Q4_K_M GGUF.

Links

Model details

Property Value
Base model nvidia/Minitron-4B-Base (merge)
Architecture Nemotron (4.2B parameters)
Format GGUF v3, Q4_K_M (4-bit)
Context length 4,096 tokens
Layers / hidden size 32 / 3,072
Attention 24 heads, 8 KV heads (GQA)
Feed-forward size 9,216
Vocabulary 256,000 tokens
RoPE base 10000, 64 dims
Chat template Nemotron (<extra_id_0>System, <extra_id_1>User, <extra_id_1>Assistant)
File snapkitty-merged.Q4_K_M.gguf, 2.71 GB
SHA-256 4d959da2985affc363c02784307ac135a69e77768d8d5475c6f15bb99560c59f

Run it locally

Ollama

ollama run hf.co/Snapkitty/snapkitty-merged:Q4_K_M

llama.cpp

llama-cli -hf Snapkitty/snapkitty-merged:Q4_K_M -cnv
# or a local file
llama-server -m snapkitty-merged.Q4_K_M.gguf -c 4096

LM Studio: search for Snapkitty/snapkitty-merged and pick the Q4_K_M file.

Python (llama-cpp-python)

from llama_cpp import Llama
llm = Llama.from_pretrained(repo_id="Snapkitty/snapkitty-merged", filename="snapkitty-merged.Q4_K_M.gguf", n_ctx=4096)
out = llm.create_chat_completion(messages=[{"role": "user", "content": "Explain SUBLEQ in one paragraph."}])
print(out["choices"][0]["message"]["content"])

Verify the download

sha256sum snapkitty-merged.Q4_K_M.gguf
# 4d959da2985affc363c02784307ac135a69e77768d8d5475c6f15bb99560c59f

Citation

@misc{snapkitty_snapkitty_merged_2026,
  title        = {snapkitty-merged: Nemotron 4B GGUF (Q4\_K\_M)},
  author       = {{Snapkitty Collective LLC}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Snapkitty/snapkitty-merged}}
}

Please also cite the base model, NVIDIA Minitron-4B-Base.

License

The model weights are a derivative of NVIDIA Minitron 4B / Nemotron-Mini-4B models and are distributed under the NVIDIA Community Model License. See LICENSE.

Downloads last month
69
GGUF
Model size
4B params
Architecture
nemotron
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Snapkitty/snapkitty-merged

Merge model
this model

Space using Snapkitty/snapkitty-merged 1

Collection including Snapkitty/snapkitty-merged