File size: 4,456 Bytes
f0d2311
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
---
license: apache-2.0
base_model:
- google/gemma-4-E4B-it
tags:
- gemma
- gemma-4
- gguf
- multimodal
- vision
- finetune
- sft
- russian
- physics
- math
- coding
library_name: gguf
pipeline_tag: text-generation
---

<div align="center">

# gemma-cvantic

**multimodal SFT fine-tune of Google's `gemma-4-E4B-it` · GGUF**

[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-gemma--cvantic-FF9D00)](https://huggingface.co/devoffeed/gemma-cvantic)
[![Base](https://img.shields.io/badge/base-google%2Fgemma--4--E4B--it-blue)](https://huggingface.co/google/gemma-4-E4B-it)
[![License](https://img.shields.io/badge/license-Apache--2.0-green)](https://www.apache.org/licenses/LICENSE-2.0)
[![GGUF](https://img.shields.io/badge/format-GGUF%20%2F%20llama.cpp-yellow)](https://github.com/ggml-org/llama.cpp)
[![Quant: Q8/Q5/iQ4](https://img.shields.io/badge/quants-Q8_0%20%E2%80%A2%20Q5_K_S%20%E2%80%A2%20iQ4_XS-orange)]()

</div>

---

## What is this?

An experimental **instruction fine-tune** of [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) — the 4.5B-effective-parameter omni-modal Gemma 4 model (128K context, text + image + audio input).

Trained with **SFT (13,000 examples)** to strengthen math, coding, physics/astronomy reasoning and Russian-language instructions. Quantized to GGUF with `llama.cpp` for fast local inference.

---

## Quantization lineup

| File | Quant | Size | Quality / speed |
|---|---|---|---|
| `gemma-cvantic.Q8_0.gguf` | **Q8_0** | ~7.6 GB | near-lossless, best quality |
| `gemma-cvantic.Q5_K_S.gguf` | **Q5_K_S** | ~5.4 GB | great balance ★ recommended |
| `gemma-cvantic.IQ4_XS.gguf` | **iQ4_XS** (imatrix) | ~4.8 GB | smallest, fast, slightly less accurate |
| `gemma-cvantic.BF16-mmproj.gguf` | vision projector | ~0.9 GB | required for image input |

> `mmproj` is the **multimodal (vision) projector** — pass it with `--mmproj` to enable image understanding. All quants are BF16/FP16 conversions of the same merged weights, so any main-file + `mmproj` combo works.

---

## Training recipe

| Parameter | Value |
|---|---|
| Base model | `google/gemma-4-E4B-it` |
| Method | SFT (DoRA / LoRA-style adapter, then merged) |
| Trainable params | ~36.7M (0.61%) |
| Examples | 13,000 |
| Context | 128K (inherited) |

**Dataset mix** (MIT / Apache-2.0 only):

| Dataset | Split | Rows | Domain |
|---|---|---|---|
| `HuggingFaceH4/ultrachat_200k` | train_sft | 5,000 | general chat |
| `theblackcat102/evol-codealpaca-v1` | train | 3,000 | coding |
| `qwedsacf/competition_math` (MATH) | train | 2,000 | math |
| `HuggingFaceTB/cosmopedia` | openstax · physics/astronomy | 2,000 | physics & astronomy |
| `openai/gsm8k` | main/train | 1,000 | grade-school math |

---

## Quick start

### llama.cpp (text)

```bash
llama-cli \
  -m gemma-cvantic.Q5_K_S.gguf \
  -p "Реши задачу: если цена товара выросла на 20% и составила 480 руб., какой была исходная цена?"
```

### Multimodal (vision)

```bash
llama-cli \
  -m gemma-cvantic.Q5_K_S.gguf \
  --mmproj gemma-cvantic.BF16-mmproj.gguf \
  -i
```

```
> what's in this photo?
```

### llama-server (OpenAI-compatible API)

```bash
llama-server \
  -m gemma-cvantic.Q8_0.gguf \
  --mmproj gemma-cvantic.BF16-mmproj.gguf \
  --port 8080
```

```python
import openai

client = openai.OpenAI(base_url="http://localhost:8080/v1", api_key="local")
resp = client.chat.completions.create(
    model="gemma-cvantic",
    messages=[{"role": "user", "content": "Расскажи про эффект Доплера на пальцах"}],
)
print(resp.choices[0].message.content)
```

---

## Picking a file

- **CPU-only, want quality** → `Q5_K_S` (fits ~8 GB RAM/VRAM, sweet spot)
- **16 GB+ / strong GPU** → `Q8_0`
- **Tiny footprint / speed first** → `iQ4_XS` (fits ~6 GB)
- **Vision** → always add the `BF16-mmproj`

Sizes are approximate; VRAM usage depends on context length.

---

## Notes

- Working-name **"cvantic"** — experimental build, quality varies per domain. Think of it as a field test of Gemma 4 E4B fine-tuning.
- Text-only SFT; vision/audio behavior inherited from the base model and *not* specifically tuned.
- Base model: [Google DeepMind](https://deepmind.google/models/gemma/) · License: **Apache 2.0** · [Gemma 4 docs](https://ai.google.dev/gemma/docs/core)

---

*Quantized with `llama.cpp` (BF16 base + imatrix for iQ4_XS).*