gemma-cvantic / README.md
devoffeed's picture
Upload folder using huggingface_hub
f0d2311 verified
|
Raw History Blame Contribute Delete
4.46 kB
---
license: apache-2.0
base_model:
- google/gemma-4-E4B-it
tags:
- gemma
- gemma-4
- gguf
- multimodal
- vision
- finetune
- sft
- russian
- physics
- math
- coding
library_name: gguf
pipeline_tag: text-generation
---
<div align="center">
# gemma-cvantic
**multimodal SFT fine-tune of Google's `gemma-4-E4B-it` · GGUF**
[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-gemma--cvantic-FF9D00)](https://huggingface.co/devoffeed/gemma-cvantic)
[![Base](https://img.shields.io/badge/base-google%2Fgemma--4--E4B--it-blue)](https://huggingface.co/google/gemma-4-E4B-it)
[![License](https://img.shields.io/badge/license-Apache--2.0-green)](https://www.apache.org/licenses/LICENSE-2.0)
[![GGUF](https://img.shields.io/badge/format-GGUF%20%2F%20llama.cpp-yellow)](https://github.com/ggml-org/llama.cpp)
[![Quant: Q8/Q5/iQ4](https://img.shields.io/badge/quants-Q8_0%20%E2%80%A2%20Q5_K_S%20%E2%80%A2%20iQ4_XS-orange)]()
</div>
---
## What is this?
An experimental **instruction fine-tune** of [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) — the 4.5B-effective-parameter omni-modal Gemma 4 model (128K context, text + image + audio input).
Trained with **SFT (13,000 examples)** to strengthen math, coding, physics/astronomy reasoning and Russian-language instructions. Quantized to GGUF with `llama.cpp` for fast local inference.
---
## Quantization lineup
| File | Quant | Size | Quality / speed |
|---|---|---|---|
| `gemma-cvantic.Q8_0.gguf` | **Q8_0** | ~7.6 GB | near-lossless, best quality |
| `gemma-cvantic.Q5_K_S.gguf` | **Q5_K_S** | ~5.4 GB | great balance ★ recommended |
| `gemma-cvantic.IQ4_XS.gguf` | **iQ4_XS** (imatrix) | ~4.8 GB | smallest, fast, slightly less accurate |
| `gemma-cvantic.BF16-mmproj.gguf` | vision projector | ~0.9 GB | required for image input |
> `mmproj` is the **multimodal (vision) projector** — pass it with `--mmproj` to enable image understanding. All quants are BF16/FP16 conversions of the same merged weights, so any main-file + `mmproj` combo works.
---
## Training recipe
| Parameter | Value |
|---|---|
| Base model | `google/gemma-4-E4B-it` |
| Method | SFT (DoRA / LoRA-style adapter, then merged) |
| Trainable params | ~36.7M (0.61%) |
| Examples | 13,000 |
| Context | 128K (inherited) |
**Dataset mix** (MIT / Apache-2.0 only):
| Dataset | Split | Rows | Domain |
|---|---|---|---|
| `HuggingFaceH4/ultrachat_200k` | train_sft | 5,000 | general chat |
| `theblackcat102/evol-codealpaca-v1` | train | 3,000 | coding |
| `qwedsacf/competition_math` (MATH) | train | 2,000 | math |
| `HuggingFaceTB/cosmopedia` | openstax · physics/astronomy | 2,000 | physics & astronomy |
| `openai/gsm8k` | main/train | 1,000 | grade-school math |
---
## Quick start
### llama.cpp (text)
```bash
llama-cli \
-m gemma-cvantic.Q5_K_S.gguf \
-p "Реши задачу: если цена товара выросла на 20% и составила 480 руб., какой была исходная цена?"
```
### Multimodal (vision)
```bash
llama-cli \
-m gemma-cvantic.Q5_K_S.gguf \
--mmproj gemma-cvantic.BF16-mmproj.gguf \
-i
```
```
> what's in this photo?
```
### llama-server (OpenAI-compatible API)
```bash
llama-server \
-m gemma-cvantic.Q8_0.gguf \
--mmproj gemma-cvantic.BF16-mmproj.gguf \
--port 8080
```
```python
import openai
client = openai.OpenAI(base_url="http://localhost:8080/v1", api_key="local")
resp = client.chat.completions.create(
model="gemma-cvantic",
messages=[{"role": "user", "content": "Расскажи про эффект Доплера на пальцах"}],
)
print(resp.choices[0].message.content)
```
---
## Picking a file
- **CPU-only, want quality** → `Q5_K_S` (fits ~8 GB RAM/VRAM, sweet spot)
- **16 GB+ / strong GPU** → `Q8_0`
- **Tiny footprint / speed first** → `iQ4_XS` (fits ~6 GB)
- **Vision** → always add the `BF16-mmproj`
Sizes are approximate; VRAM usage depends on context length.
---
## Notes
- Working-name **"cvantic"** — experimental build, quality varies per domain. Think of it as a field test of Gemma 4 E4B fine-tuning.
- Text-only SFT; vision/audio behavior inherited from the base model and *not* specifically tuned.
- Base model: [Google DeepMind](https://deepmind.google/models/gemma/) · License: **Apache 2.0** · [Gemma 4 docs](https://ai.google.dev/gemma/docs/core)
---
*Quantized with `llama.cpp` (BF16 base + imatrix for iQ4_XS).*