Gemma 4 text encoders for ComfyUI (BF16 + INT8 ConvRot + W4A8)
Official Gemma 4 E4B-it, 12B-it and 31B-it text encoders for ComfyUI: E4B and 12B in BF16 and INT8 ConvRot, 31B in BF16, INT8 ConvRot and W4A8.
This repository collects ComfyUI single-file Gemma 4 instruction-tuned models, each in BF16 and INT8 ConvRot, and 31B also in W4A8, under one naming scheme: <original model name>_<bf16|int8_convrot|w4a8>.safetensors.
Three files are made here: gemma-4-31B-it_bf16, Google's shards merged into the ComfyUI single-file layout, and its two conversions, gemma-4-31B-it_int8_convrot and gemma-4-31B-it_w4a8. The other four are byte-identical re-uploads of files published by Comfy-Org, renamed only. Every source is licensed under Apache-2.0, and all credit belongs to the original authors listed below.
Usage in ComfyUI
- Put the file in
ComfyUI/models/text_encoders/. - Load it with the Load CLIP node (
CLIPLoader) and pick the type your workflow uses. ComfyUI detects the architecture (Gemma 4 E4B, 12B or 31B) from the weights. - The
_int8_convrotfiles need a current ComfyUI with comfy-kitchen INT8 ConvRot support. - The
_w4a8file needs comfy-kitchen W4A8 support and a GPU with compute capability 8.0 or newer. - The three 31B files need the load-time patch shipped with our prompter node (link to follow) until ComfyUI supports Gemma 4 31B's global-attention layout; see Gemma 4 31B.
| File | Original model | Made by | Size (bytes) | sha256 |
|---|---|---|---|---|
gemma-4-E4B-it_bf16.safetensors |
gemma-4-E4B-it | Comfy-Org (re-upload) | 16,024,746,334 | afe21e7c99d5a2ba52bc246a464d2458726204c3ce98ee81398204786ecab5ab |
gemma-4-E4B-it_int8_convrot.safetensors |
gemma-4-E4B-it | Comfy-Org (re-upload) | 8,090,965,702 | 974d0c838ef4ac1a989b06ccb4e57691c21b6270dd8e345ffa7531f9388f117c |
gemma-4-12B-it_bf16.safetensors |
gemma-4-12B-it | Comfy-Org (re-upload) | 23,951,709,034 | c00e57c007a54d8dab9b2c8b147808aa88207fcc8d61db713de009712e288199 |
gemma-4-12B-it_int8_convrot.safetensors |
gemma-4-12B-it | Comfy-Org (re-upload) | 12,055,234,634 | bf77dc0b435c487a638909d8f2ccf5a7e4c9838e7bc56545ea6e251a603c5793 |
gemma-4-31B-it_bf16.safetensors |
gemma-4-31B-it | merged here | 62,578,493,650 | 3f4fc52e8cdff5a44520b61e14921aa0e9803c1d1f06026db6c896e22b9416fa |
gemma-4-31B-it_int8_convrot.safetensors |
gemma-4-31B-it | converted here | 31,357,238,261 | 7476779a2adf797e4ea1ff09fc175527a00e70bf0a7431c02f23745e160c2358 |
gemma-4-31B-it_w4a8.safetensors |
gemma-4-31B-it | converted here | 18,544,425,671 | 63ff306b80d63baf4bcac84f2b296d25620edd4f45f233c4700cd1254508c42d |
Every re-uploaded file has the same sha256 as the LFS object in its source repository.
Gemma 4 E4B
gemma-4-E4B-it_bf16.safetensors
- Role: the official Gemma 4 E4B-it in BF16, in the ComfyUI single-file layout.
- Source: Comfy-Org/gemma-4, original filename
text_encoders/gemma4_e4b_it_bf16.safetensors, revision63d0f7c476756b88910170c1df75e2384ea1af31. It is Comfy-Org's ComfyUI packaging of google/gemma-4-E4B-it by Google DeepMind. - License: Apache-2.0, as stated by Google (Gemma 4 license).
- Changes: none. Byte-identical, renamed.
gemma-4-E4B-it_int8_convrot.safetensors
- Role: Comfy-Org's official INT8 ConvRot build of Gemma 4 E4B-it.
- Source: Comfy-Org/gemma-4, original filename
text_encoders/gemma4_e4b_it_int8_convrot.safetensors, revision63d0f7c476756b88910170c1df75e2384ea1af31. Model by Google DeepMind, quantization by Comfy-Org. - License: Apache-2.0, as stated by Google (Gemma 4 license).
- Changes: none. Byte-identical, renamed.
Gemma 4 12B
gemma-4-12B-it_bf16.safetensors
- Role: the official Gemma 4 12B-it in BF16, in the ComfyUI single-file layout.
- Source: Comfy-Org/gemma-4, original filename
text_encoders/gemma4_12b_bf16.safetensors, revision63d0f7c476756b88910170c1df75e2384ea1af31. It is Comfy-Org's ComfyUI packaging of google/gemma-4-12B-it by Google DeepMind. Comfy-Org's filename has no_it; their model card listsgoogle/gemma-4-12B-itas the base model. - License: Apache-2.0, as stated by Google (Gemma 4 license).
- Changes: none. Byte-identical, renamed.
gemma-4-12B-it_int8_convrot.safetensors
- Role: Comfy-Org's official INT8 ConvRot build of Gemma 4 12B-it.
- Source: Comfy-Org/gemma-4, original filename
text_encoders/gemma4_12b_int8_convrot.safetensors, revision63d0f7c476756b88910170c1df75e2384ea1af31. Model by Google DeepMind, quantization by Comfy-Org. - License: Apache-2.0, as stated by Google (Gemma 4 license).
- Changes: none. Byte-identical, renamed.
Gemma 4 31B
Gemma 4 31B's 10 global-attention layers use attention_k_eq_v: 4 key/value heads of size 512 and no v_proj, the key projection also serves as the value. ComfyUI's Gemma4_31B_Config (checked at ComfyUI commit a7169322) expects 16 KV heads and a v_proj on those layers. The three files below keep Google's layout unchanged. Until ComfyUI supports it, load them with our prompter node (link to follow) installed: the load-time patch shipped with it adapts ComfyUI's Gemma 4 31B config. The patch is part of the node itself; there is no separate patch node to install.
gemma-4-31B-it_bf16.safetensors
- Role: the official Gemma 4 31B-it in BF16 as one file in the ComfyUI key layout. It is also the source for
gemma-4-31B-it_int8_convrotandgemma-4-31B-it_w4a8. - Source: the 2 safetensors shards of google/gemma-4-31B-it at revision
842da3794eaa0b77d5f08bae87a17459d91ff475, merged here into one file. Model by Google DeepMind. - License: Apache-2.0, as stated by Google (Gemma 4 license).
- Changes:
- the shards merged into one file by a streaming copy;
- keys renamed to the layout of Comfy-Org's Gemma 4 E4B and 12B files:
model.language_model.โmodel.,model.vision_tower.โvision_model.,model.embed_vision.โmulti_modal_projector.; - a U8 tensor
tokenizer_jsonadded, holding the model'stokenizer.json(sha256cc8d3a0ce36466ccc1278bf987df5f71db1719b9ca6b4118264f45cb627bfe0f), which ComfyUI's Gemma 4 tokenizer reads from the file; - tensor dtypes, shapes and bytes otherwise unchanged; the global-attention layers keep Google's layout (see above).
- Loading: needs the load-time patch shipped with our prompter node (see above).
gemma-4-31B-it_int8_convrot.safetensors
- Role: INT8 ConvRot build of the official Gemma 4 31B-it. It is the quality reference, the quantized build closest to BF16, and needs offload on a 24 GB card.
- Source: our conversion of
gemma-4-31B-it_bf16.safetensorsabove (its name and sha256 are recorded in the file's__metadata__). The base model is by Google DeepMind. - License: Apache-2.0, as stated by Google (Gemma 4 license).
- Changes: INT8 ConvRot quantization (recipe below); the other tensors are passed through in their source dtype.
- Recipe: Comfy-Org's
quant_int8_convrot.pyfrom comfy-model-tools at commit1846ff1a, without arguments, and its own layer selection: the block linears and the token embedding become INT8 ConvRot (FP32 rotation, per-row absmax scale), everything else is passed through. - Quantized layers: 600 (
.comfy_quantentries in the file). - Quantizer check: the quantizer's own reconstruction check on its 599 ConvRot block linears (dequantized with the inverse rotation, compared with the BF16 source): cosine min 0.99991, relative error mean 0.89 % / max 1.30 %.
- Layout (as published): the 410 language-model linears,
embed_tokensand the 189 vision-tower linears are INT8 ConvRot. The global-attention layers keep Google's layout (see above). - Loading: needs the load-time patch shipped with our prompter node (see above).
gemma-4-31B-it_w4a8.safetensors
- Role: W4A8 build of the official Gemma 4 31B-it.
- Source: our conversion of
gemma-4-31B-it_bf16.safetensorsabove (its name and sha256 are recorded in the file's__metadata__). The base model is by Google DeepMind. - License: Apache-2.0, as stated by Google (Gemma 4 license).
- Changes: W4A8 quantization (recipe below); the other tensors are passed through in their source dtype.
- Recipe: Comfy-Org's
quant_int8_convrot.pyfrom comfy-model-tools at commit1846ff1awith--w4a8, and its own layer selection: a selected block linear whose input dimension is a multiple of 256 and that has at least 64 rows becomesasym_w4a8_int8(4-bit weights with a codebook and FP8 scales per group of 16, ConvRot 256; activations quantized to INT8 at run time), the other selected linears and the token embedding become INT8 ConvRot, everything else is passed through. - Quantized layers: 600 (
.comfy_quantentries in the file). - Quantizer check: the quantizer's own reconstruction check, dequantized and compared with the BF16 source: its 410
asym_w4a8_int8linears relative error mean 7.32 % / max 7.35 % (the script computes no cosine for them); its 189 INT8 ConvRot linears cosine min 0.99995, relative error mean 0.82 % / max 1.03 %. - Layout (as published): the 410 language-model linears are
asym_w4a8_int8(group size 16, ConvRot 256);embed_tokensand the 189 vision-tower linears are INT8 ConvRot. The global-attention layers keep Google's layout (see above). - Requirements: comfy-kitchen's W4A8 kernels, which need a GPU with compute capability 8.0 or newer.
- Loading: needs the load-time patch shipped with our prompter node (see above).