Instructions to use ZichenAI/Baihu-V1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ZichenAI/Baihu-V1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ZichenAI/Baihu-V1-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ZichenAI/Baihu-V1-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZichenAI/Baihu-V1-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZichenAI/Baihu-V1-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ZichenAI/Baihu-V1-GGUF:Q4_K_M
- Ollama
How to use ZichenAI/Baihu-V1-GGUF with Ollama:
ollama run hf.co/ZichenAI/Baihu-V1-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ZichenAI/Baihu-V1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ZichenAI/Baihu-V1-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ZichenAI/Baihu-V1-GGUF with Docker Model Runner:
docker model run hf.co/ZichenAI/Baihu-V1-GGUF:Q4_K_M
- Lemonade
How to use ZichenAI/Baihu-V1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ZichenAI/Baihu-V1-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Baihu-V1-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ZichenAI/Baihu-V1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ZichenAI/Baihu-V1-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ZichenAI/Baihu-V1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ZichenAI/Baihu-V1-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Baihu-V1 (GGUF)
- Files
- Quick start
- Benchmarks: base vs Baihu-V1
- 1. Held-out transfer-dialogue benchmark — 321 conversations, 12 topic groups never seen in training
- 2. Same benchmark, measured on the shipped GGUF (Q4_K_M) via
llama-server - 3. General capability — MMLU, 5-shot, option log-likelihood
- 4. Held-out perplexity (
llama-perplexity, Q4_K_M, 60 × 2048-token chunks) - 5. Inference speed (
llama-bench, RTX 4090 D, Q4_K_M,-ngl 99)
- Training details
- Intended use and limitations
- License
- Citation
- Files
Baihu-V1 (GGUF)
Baihu-V1 is a LoRA fine-tune of google/gemma-4-E2B-it on a curated bilingual
(Chinese + English) corpus of multi-turn transfer-dialogue data: conversations
where a rule, analogy or format is introduced once and then carried to new cases
across several turns.
The fine-tune changes the model's answering behaviour rather than its knowledge. It replies in short, plain prose, states the result directly, and stops, instead of producing long markdown lecture-style answers. Held-out evaluation shows this costs essentially nothing on general knowledge (MMLU 5-shot 53.4% → 52.6%) while roughly doubling reply quality against the reference style (ROUGE-L 0.236 → 0.432).
Files
| File | Format | Size | Notes |
|---|---|---|---|
Baihu-V1-Q4_K_M.gguf |
GGUF Q4_K_M | 3.42 GB | Recommended |
Baihu-V1-Q8_0.gguf |
GGUF Q8_0 | 4.95 GB | Near-lossless |
Both files are text-only conversions of the merged checkpoint. The vision and audio towers of Gemma 4 E2B are not used for text generation and were not fine-tuned, so they are not included.
Quick start
# llama.cpp CLI
llama-cli -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M \
-p "<bos><|turn>user\nIn one sentence: why does metal feel colder than wood?<turn|>\n<|turn>model\n" \
-ngl 99 -st
# llama-server (OpenAI-compatible)
llama-server -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M -c 8192 -ngl 99
The model uses the standard Gemma 4 chat template, so any client that speaks the
Gemma 4 format works (llama.cpp, LM Studio, Ollama, transformers).
Benchmarks: base vs Baihu-V1
Every number below comes from the same scripts, the same held-out split and the same
hardware (1× RTX 4090 D), so the delta is what matters, not the absolute values.
The raw result files are in benchmarks/ and a consolidated copy is in
benchmarks/summary.json.
1. Held-out transfer-dialogue benchmark — 321 conversations, 12 topic groups never seen in training
Each conversation is truncated before its final assistant turn; the model must
produce that turn. ROUGE-L is character-level LCS F1 against the reference reply.
"Plain text" means the reply contains no markdown markup at all.
Len ratio is reply length ÷ reference length (1.00 = matching the reference style).
| Metric | gemma-4-E2B-it (base) | Baihu-V1 | Δ |
|---|---|---|---|
| ROUGE-L ↑ | 0.236 | 0.432 | +0.196 |
| Plain-text rate ↑ | 1.6 % | 100 % | +98.4 pp |
| Reply length ratio (target 1.00) | 3.86 | 0.97 | −2.89 |
| Mean reply length (chars) | 432.3 | 113.7 | −318.6 |
| Key-number recall ↑ | 0.580 | 0.528 | −0.052 |
On the key-number recall row: the metric asks what fraction of the reference's numbers appear in the reply, which rewards saying more. The base model emits ~4× more text and therefore gets more chances to hit a number; the fine-tune trades a little of that for conciseness, while ROUGE-L — which also penalises the extra text — nearly doubles. No individual subject or pattern loses more than 6 pp.
2. Same benchmark, measured on the shipped GGUF (Q4_K_M) via llama-server
This row set is the end-to-end check of the published artefacts themselves.
| Metric | gemma-4-E2B-it-Q4_K_M | Baihu-V1-Q4_K_M | Δ |
|---|---|---|---|
| ROUGE-L ↑ | 0.235 | 0.419 | +0.184 |
| Plain-text rate ↑ | 1.3 % | 100 % | +98.7 pp |
| Reply length ratio (target 1.00) | 3.85 | 0.96 | −2.89 |
| Mean reply length (chars) | 432.2 | 113.9 | −318.3 |
| Wall-clock for 321 replies ↓ | 192.1 s | 62.1 s | −67.7 % |
The GGUF reproduces the merged HF checkpoint closely (ROUGE-L 0.419 vs 0.432), which confirms that the merge-and-convert pipeline and the Q4_K_M quantisation introduce no meaningful degradation.
3. General capability — MMLU, 5-shot, option log-likelihood
500 test questions from 10 subjects (cais/mmlu, 50 per subject, 5-shot dev prompts).
Scored by the log-likelihood of the option letter, the same protocol llama.cpp uses.
| Subject | gemma-4-E2B-it | Baihu-V1 | Δ |
|---|---|---|---|
| astronomy | 0.540 | 0.520 | −0.020 |
| computer_security | 0.440 | 0.500 | +0.060 |
| elementary_mathematics | 0.480 | 0.480 | 0.000 |
| formal_logic | 0.580 | 0.600 | +0.020 |
| high_school_biology | 0.540 | 0.540 | 0.000 |
| high_school_mathematics | 0.440 | 0.400 | −0.040 |
| high_school_physics | 0.380 | 0.480 | +0.100 |
| high_school_world_history | 0.640 | 0.640 | 0.000 |
| logical_fallacies | 0.640 | 0.540 | −0.100 |
| world_religions | 0.660 | 0.560 | −0.100 |
| Overall | 0.534 | 0.526 | −0.008 |
The −0.8 pp shift is inside the sampling noise of a 500-question suite (≈ ±2.2 pp at 95 % confidence), so the fine-tune did not cause measurable forgetting. This is not a general-purpose knowledge upgrade — it is a behaviour and style tune on top of the same base.
4. Held-out perplexity (llama-perplexity, Q4_K_M, 60 × 2048-token chunks)
| Model | Perplexity ↓ |
|---|---|
| gemma-4-E2B-it-Q4_K_M | 13.74 |
| Baihu-V1-Q4_K_M | 5.28 |
5. Inference speed (llama-bench, RTX 4090 D, Q4_K_M, -ngl 99)
| Model | Prefill (t/s) | Decode (t/s) |
|---|---|---|
| gemma-4-E2B-it-Q4_K_M | 15 095 | 307.8 |
| Baihu-V1-Q4_K_M | 15 643 | 307.1 |
Fine-tuning a subset of the attention and MLP projections does not change the architecture, so throughput matches the base model within run-to-run noise.
Training details
| Base model | google/gemma-4-E2B-it (5.13 B total params) |
| Method | LoRA, rank 16, alpha 32, dropout 0.05, bias none |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj inside model.language_model (205 modules) |
| Trainable params | 24,158,208 / 5,128,455,712 (0.47 %) |
| Train / val | 973 / 321 conversations, split by topic group so validation topics are unseen |
| Loss | assistant turns only (user turns masked with −100) |
| Epochs / steps | 3 / 183 optimiser steps |
| Batch | 4 × gradient accumulation 4 (effective 16) |
| LR | 2e-4, cosine, 5 warm-up steps, max grad norm 1.0 |
| Precision / context | bf16 / 768 tokens (train lengths: min 155 · p50 290 · p90 385 · max 545) |
| Optimiser | AdamW (PyTorch) |
| Hardware / time | 1× RTX 4090 D (24 GB) / 6.2 minutes |
| Final loss | train 1.165 · held-out eval 1.536 |
| Framework | transformers 5.17.0, peft 0.21.1, torch 2.14.0+cu130 |
Data
973 training / 321 validation multi-turn conversations (6–10 messages each),
synthesised around six transfer patterns — rule_induction, analogy_transfer,
format_transfer, counterfactual, cross_domain, teaching_loop — in both
Chinese and English. The validation split holds out whole topic groups, so every
validation conversation is about a topic the model never saw during training.
Export pipeline
- Merge the LoRA adapter into the base weights in bf16 (
peft.merge_and_unload). convert_hf_to_gguf.py --outtype bf16(llama.cpp).llama-quantize … Q4_K_M/Q8_0.- Re-measure everything on the resulting GGUF with
llama-serverandllama-perplexity.
Intended use and limitations
Use it for: concise, direct question answering in Chinese or English; multi-turn explanatory dialogues where a short answer is preferred over a formatted essay; on-device or edge deployment where a ~3.4 GB 2B-class model is the budget.
Do not rely on it for: factual lookups, mathematics beyond short arithmetic, long-form writing, tool calling, or anything needing the vision/audio capabilities of the base Gemma 4 E2B — those towers are not part of this GGUF.
The model was fine-tuned on synthetic dialogue data. It can still hallucinate, it
inherits the biases of its base model, and its output should be reviewed before being
used in any consequential setting. It ships at Q4_K_M by default — use the Q8_0
file when near-lossless output matters.
License
Custom license — personal use free, commercial use requires a paid licence.
| Use | Cost |
|---|---|
| Personal, research, academic, non-profit, open-source non-commercial | Free |
| Any commercial or revenue-generating use | Paid licence required |
Commercial licensing contact:
novaweb6868@outlook.com Suggested subject:
Baihu-V1 Commercial License
See LICENSE for the full text. This model is a derivative of
google/gemma-4-E2B-it and remains subject to the
Gemma 4 license terms. Where the
two conflict, the Gemma 4 terms prevail.
Citation
@misc{baihu-v1,
title = {Baihu-V1: a transfer-dialogue LoRA fine-tune of Gemma 4 E2B, exported to GGUF},
author = {ZichenAI},
year = {2026},
url = {https://huggingface.co/ZichenAI/Baihu-V1-GGUF}
}
- Downloads last month
- 173
4-bit
8-bit