Baihu-V1 (GGUF)

Baihu-V1 is a LoRA fine-tune of google/gemma-4-E2B-it on a curated bilingual (Chinese + English) corpus of multi-turn transfer-dialogue data: conversations where a rule, analogy or format is introduced once and then carried to new cases across several turns.

The fine-tune changes the model's answering behaviour rather than its knowledge. It replies in short, plain prose, states the result directly, and stops, instead of producing long markdown lecture-style answers. Held-out evaluation shows this costs essentially nothing on general knowledge (MMLU 5-shot 53.4% → 52.6%) while roughly doubling reply quality against the reference style (ROUGE-L 0.236 → 0.432).


Files

File Format Size Notes
Baihu-V1-Q4_K_M.gguf GGUF Q4_K_M 3.42 GB Recommended
Baihu-V1-Q8_0.gguf GGUF Q8_0 4.95 GB Near-lossless

Both files are text-only conversions of the merged checkpoint. The vision and audio towers of Gemma 4 E2B are not used for text generation and were not fine-tuned, so they are not included.

Quick start

# llama.cpp CLI
llama-cli -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M \
  -p "<bos><|turn>user\nIn one sentence: why does metal feel colder than wood?<turn|>\n<|turn>model\n" \
  -ngl 99 -st

# llama-server (OpenAI-compatible)
llama-server -hf ZichenAI/Baihu-V1-GGUF:Q4_K_M -c 8192 -ngl 99

The model uses the standard Gemma 4 chat template, so any client that speaks the Gemma 4 format works (llama.cpp, LM Studio, Ollama, transformers).


Benchmarks: base vs Baihu-V1

Every number below comes from the same scripts, the same held-out split and the same hardware (1× RTX 4090 D), so the delta is what matters, not the absolute values. The raw result files are in benchmarks/ and a consolidated copy is in benchmarks/summary.json.

1. Held-out transfer-dialogue benchmark — 321 conversations, 12 topic groups never seen in training

Each conversation is truncated before its final assistant turn; the model must produce that turn. ROUGE-L is character-level LCS F1 against the reference reply. "Plain text" means the reply contains no markdown markup at all. Len ratio is reply length ÷ reference length (1.00 = matching the reference style).

Metric gemma-4-E2B-it (base) Baihu-V1 Δ
ROUGE-L ↑ 0.236 0.432 +0.196
Plain-text rate ↑ 1.6 % 100 % +98.4 pp
Reply length ratio (target 1.00) 3.86 0.97 −2.89
Mean reply length (chars) 432.3 113.7 −318.6
Key-number recall ↑ 0.580 0.528 −0.052

On the key-number recall row: the metric asks what fraction of the reference's numbers appear in the reply, which rewards saying more. The base model emits ~4× more text and therefore gets more chances to hit a number; the fine-tune trades a little of that for conciseness, while ROUGE-L — which also penalises the extra text — nearly doubles. No individual subject or pattern loses more than 6 pp.

2. Same benchmark, measured on the shipped GGUF (Q4_K_M) via llama-server

This row set is the end-to-end check of the published artefacts themselves.

Metric gemma-4-E2B-it-Q4_K_M Baihu-V1-Q4_K_M Δ
ROUGE-L ↑ 0.235 0.419 +0.184
Plain-text rate ↑ 1.3 % 100 % +98.7 pp
Reply length ratio (target 1.00) 3.85 0.96 −2.89
Mean reply length (chars) 432.2 113.9 −318.3
Wall-clock for 321 replies ↓ 192.1 s 62.1 s −67.7 %

The GGUF reproduces the merged HF checkpoint closely (ROUGE-L 0.419 vs 0.432), which confirms that the merge-and-convert pipeline and the Q4_K_M quantisation introduce no meaningful degradation.

3. General capability — MMLU, 5-shot, option log-likelihood

500 test questions from 10 subjects (cais/mmlu, 50 per subject, 5-shot dev prompts). Scored by the log-likelihood of the option letter, the same protocol llama.cpp uses.

Subject gemma-4-E2B-it Baihu-V1 Δ
astronomy 0.540 0.520 −0.020
computer_security 0.440 0.500 +0.060
elementary_mathematics 0.480 0.480 0.000
formal_logic 0.580 0.600 +0.020
high_school_biology 0.540 0.540 0.000
high_school_mathematics 0.440 0.400 −0.040
high_school_physics 0.380 0.480 +0.100
high_school_world_history 0.640 0.640 0.000
logical_fallacies 0.640 0.540 −0.100
world_religions 0.660 0.560 −0.100
Overall 0.534 0.526 −0.008

The −0.8 pp shift is inside the sampling noise of a 500-question suite (≈ ±2.2 pp at 95 % confidence), so the fine-tune did not cause measurable forgetting. This is not a general-purpose knowledge upgrade — it is a behaviour and style tune on top of the same base.

4. Held-out perplexity (llama-perplexity, Q4_K_M, 60 × 2048-token chunks)

Model Perplexity ↓
gemma-4-E2B-it-Q4_K_M 13.74
Baihu-V1-Q4_K_M 5.28

5. Inference speed (llama-bench, RTX 4090 D, Q4_K_M, -ngl 99)

Model Prefill (t/s) Decode (t/s)
gemma-4-E2B-it-Q4_K_M 15 095 307.8
Baihu-V1-Q4_K_M 15 643 307.1

Fine-tuning a subset of the attention and MLP projections does not change the architecture, so throughput matches the base model within run-to-run noise.


Training details

Base model google/gemma-4-E2B-it (5.13 B total params)
Method LoRA, rank 16, alpha 32, dropout 0.05, bias none
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj inside model.language_model (205 modules)
Trainable params 24,158,208 / 5,128,455,712 (0.47 %)
Train / val 973 / 321 conversations, split by topic group so validation topics are unseen
Loss assistant turns only (user turns masked with −100)
Epochs / steps 3 / 183 optimiser steps
Batch 4 × gradient accumulation 4 (effective 16)
LR 2e-4, cosine, 5 warm-up steps, max grad norm 1.0
Precision / context bf16 / 768 tokens (train lengths: min 155 · p50 290 · p90 385 · max 545)
Optimiser AdamW (PyTorch)
Hardware / time 1× RTX 4090 D (24 GB) / 6.2 minutes
Final loss train 1.165 · held-out eval 1.536
Framework transformers 5.17.0, peft 0.21.1, torch 2.14.0+cu130

Data

973 training / 321 validation multi-turn conversations (6–10 messages each), synthesised around six transfer patterns — rule_induction, analogy_transfer, format_transfer, counterfactual, cross_domain, teaching_loop — in both Chinese and English. The validation split holds out whole topic groups, so every validation conversation is about a topic the model never saw during training.

Export pipeline

  1. Merge the LoRA adapter into the base weights in bf16 (peft.merge_and_unload).
  2. convert_hf_to_gguf.py --outtype bf16 (llama.cpp).
  3. llama-quantize … Q4_K_M / Q8_0.
  4. Re-measure everything on the resulting GGUF with llama-server and llama-perplexity.

Intended use and limitations

Use it for: concise, direct question answering in Chinese or English; multi-turn explanatory dialogues where a short answer is preferred over a formatted essay; on-device or edge deployment where a ~3.4 GB 2B-class model is the budget.

Do not rely on it for: factual lookups, mathematics beyond short arithmetic, long-form writing, tool calling, or anything needing the vision/audio capabilities of the base Gemma 4 E2B — those towers are not part of this GGUF.

The model was fine-tuned on synthetic dialogue data. It can still hallucinate, it inherits the biases of its base model, and its output should be reviewed before being used in any consequential setting. It ships at Q4_K_M by default — use the Q8_0 file when near-lossless output matters.


License

Custom license — personal use free, commercial use requires a paid licence.

Use Cost
Personal, research, academic, non-profit, open-source non-commercial Free
Any commercial or revenue-generating use Paid licence required

Commercial licensing contact:

novaweb6868@outlook.com Suggested subject: Baihu-V1 Commercial License

See LICENSE for the full text. This model is a derivative of google/gemma-4-E2B-it and remains subject to the Gemma 4 license terms. Where the two conflict, the Gemma 4 terms prevail.

Citation

@misc{baihu-v1,
  title  = {Baihu-V1: a transfer-dialogue LoRA fine-tune of Gemma 4 E2B, exported to GGUF},
  author = {ZichenAI},
  year   = {2026},
  url    = {https://huggingface.co/ZichenAI/Baihu-V1-GGUF}
}
Downloads last month
173
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZichenAI/Baihu-V1-GGUF

Finetuned
(385)
this model