Gemma 4 text encoders for ComfyUI (BF16 + INT8 ConvRot + W4A8)

Official Gemma 4 E4B-it, 12B-it and 31B-it text encoders for ComfyUI: E4B and 12B in BF16 and INT8 ConvRot, 31B in BF16, INT8 ConvRot and W4A8.

This repository collects ComfyUI single-file Gemma 4 instruction-tuned models, each in BF16 and INT8 ConvRot, and 31B also in W4A8, under one naming scheme: <original model name>_<bf16|int8_convrot|w4a8>.safetensors.

Three files are made here: gemma-4-31B-it_bf16, Google's shards merged into the ComfyUI single-file layout, and its two conversions, gemma-4-31B-it_int8_convrot and gemma-4-31B-it_w4a8. The other four are byte-identical re-uploads of files published by Comfy-Org, renamed only. Every source is licensed under Apache-2.0, and all credit belongs to the original authors listed below.

Usage in ComfyUI

  1. Put the file in ComfyUI/models/text_encoders/.
  2. Load it with the Load CLIP node (CLIPLoader) and pick the type your workflow uses. ComfyUI detects the architecture (Gemma 4 E4B, 12B or 31B) from the weights.
  3. The _int8_convrot files need a current ComfyUI with comfy-kitchen INT8 ConvRot support.
  4. The _w4a8 file needs comfy-kitchen W4A8 support and a GPU with compute capability 8.0 or newer.
  5. The three 31B files need the load-time patch shipped with our prompter node (link to follow) until ComfyUI supports Gemma 4 31B's global-attention layout; see Gemma 4 31B.
File Original model Made by Size (bytes) sha256
gemma-4-E4B-it_bf16.safetensors gemma-4-E4B-it Comfy-Org (re-upload) 16,024,746,334 afe21e7c99d5a2ba52bc246a464d2458726204c3ce98ee81398204786ecab5ab
gemma-4-E4B-it_int8_convrot.safetensors gemma-4-E4B-it Comfy-Org (re-upload) 8,090,965,702 974d0c838ef4ac1a989b06ccb4e57691c21b6270dd8e345ffa7531f9388f117c
gemma-4-12B-it_bf16.safetensors gemma-4-12B-it Comfy-Org (re-upload) 23,951,709,034 c00e57c007a54d8dab9b2c8b147808aa88207fcc8d61db713de009712e288199
gemma-4-12B-it_int8_convrot.safetensors gemma-4-12B-it Comfy-Org (re-upload) 12,055,234,634 bf77dc0b435c487a638909d8f2ccf5a7e4c9838e7bc56545ea6e251a603c5793
gemma-4-31B-it_bf16.safetensors gemma-4-31B-it merged here 62,578,493,650 3f4fc52e8cdff5a44520b61e14921aa0e9803c1d1f06026db6c896e22b9416fa
gemma-4-31B-it_int8_convrot.safetensors gemma-4-31B-it converted here 31,357,238,261 7476779a2adf797e4ea1ff09fc175527a00e70bf0a7431c02f23745e160c2358
gemma-4-31B-it_w4a8.safetensors gemma-4-31B-it converted here 18,544,425,671 63ff306b80d63baf4bcac84f2b296d25620edd4f45f233c4700cd1254508c42d

Every re-uploaded file has the same sha256 as the LFS object in its source repository.


Gemma 4 E4B

gemma-4-E4B-it_bf16.safetensors

  • Role: the official Gemma 4 E4B-it in BF16, in the ComfyUI single-file layout.
  • Source: Comfy-Org/gemma-4, original filename text_encoders/gemma4_e4b_it_bf16.safetensors, revision 63d0f7c476756b88910170c1df75e2384ea1af31. It is Comfy-Org's ComfyUI packaging of google/gemma-4-E4B-it by Google DeepMind.
  • License: Apache-2.0, as stated by Google (Gemma 4 license).
  • Changes: none. Byte-identical, renamed.

gemma-4-E4B-it_int8_convrot.safetensors

  • Role: Comfy-Org's official INT8 ConvRot build of Gemma 4 E4B-it.
  • Source: Comfy-Org/gemma-4, original filename text_encoders/gemma4_e4b_it_int8_convrot.safetensors, revision 63d0f7c476756b88910170c1df75e2384ea1af31. Model by Google DeepMind, quantization by Comfy-Org.
  • License: Apache-2.0, as stated by Google (Gemma 4 license).
  • Changes: none. Byte-identical, renamed.

Gemma 4 12B

gemma-4-12B-it_bf16.safetensors

  • Role: the official Gemma 4 12B-it in BF16, in the ComfyUI single-file layout.
  • Source: Comfy-Org/gemma-4, original filename text_encoders/gemma4_12b_bf16.safetensors, revision 63d0f7c476756b88910170c1df75e2384ea1af31. It is Comfy-Org's ComfyUI packaging of google/gemma-4-12B-it by Google DeepMind. Comfy-Org's filename has no _it; their model card lists google/gemma-4-12B-it as the base model.
  • License: Apache-2.0, as stated by Google (Gemma 4 license).
  • Changes: none. Byte-identical, renamed.

gemma-4-12B-it_int8_convrot.safetensors

  • Role: Comfy-Org's official INT8 ConvRot build of Gemma 4 12B-it.
  • Source: Comfy-Org/gemma-4, original filename text_encoders/gemma4_12b_int8_convrot.safetensors, revision 63d0f7c476756b88910170c1df75e2384ea1af31. Model by Google DeepMind, quantization by Comfy-Org.
  • License: Apache-2.0, as stated by Google (Gemma 4 license).
  • Changes: none. Byte-identical, renamed.

Gemma 4 31B

Gemma 4 31B's 10 global-attention layers use attention_k_eq_v: 4 key/value heads of size 512 and no v_proj, the key projection also serves as the value. ComfyUI's Gemma4_31B_Config (checked at ComfyUI commit a7169322) expects 16 KV heads and a v_proj on those layers. The three files below keep Google's layout unchanged. Until ComfyUI supports it, load them with our prompter node (link to follow) installed: the load-time patch shipped with it adapts ComfyUI's Gemma 4 31B config. The patch is part of the node itself; there is no separate patch node to install.

gemma-4-31B-it_bf16.safetensors

  • Role: the official Gemma 4 31B-it in BF16 as one file in the ComfyUI key layout. It is also the source for gemma-4-31B-it_int8_convrot and gemma-4-31B-it_w4a8.
  • Source: the 2 safetensors shards of google/gemma-4-31B-it at revision 842da3794eaa0b77d5f08bae87a17459d91ff475, merged here into one file. Model by Google DeepMind.
  • License: Apache-2.0, as stated by Google (Gemma 4 license).
  • Changes:
    • the shards merged into one file by a streaming copy;
    • keys renamed to the layout of Comfy-Org's Gemma 4 E4B and 12B files: model.language_model. โ†’ model., model.vision_tower. โ†’ vision_model., model.embed_vision. โ†’ multi_modal_projector.;
    • a U8 tensor tokenizer_json added, holding the model's tokenizer.json (sha256 cc8d3a0ce36466ccc1278bf987df5f71db1719b9ca6b4118264f45cb627bfe0f), which ComfyUI's Gemma 4 tokenizer reads from the file;
    • tensor dtypes, shapes and bytes otherwise unchanged; the global-attention layers keep Google's layout (see above).
  • Loading: needs the load-time patch shipped with our prompter node (see above).

gemma-4-31B-it_int8_convrot.safetensors

  • Role: INT8 ConvRot build of the official Gemma 4 31B-it. It is the quality reference, the quantized build closest to BF16, and needs offload on a 24 GB card.
  • Source: our conversion of gemma-4-31B-it_bf16.safetensors above (its name and sha256 are recorded in the file's __metadata__). The base model is by Google DeepMind.
  • License: Apache-2.0, as stated by Google (Gemma 4 license).
  • Changes: INT8 ConvRot quantization (recipe below); the other tensors are passed through in their source dtype.
  • Recipe: Comfy-Org's quant_int8_convrot.py from comfy-model-tools at commit 1846ff1a, without arguments, and its own layer selection: the block linears and the token embedding become INT8 ConvRot (FP32 rotation, per-row absmax scale), everything else is passed through.
  • Quantized layers: 600 (.comfy_quant entries in the file).
  • Quantizer check: the quantizer's own reconstruction check on its 599 ConvRot block linears (dequantized with the inverse rotation, compared with the BF16 source): cosine min 0.99991, relative error mean 0.89 % / max 1.30 %.
  • Layout (as published): the 410 language-model linears, embed_tokens and the 189 vision-tower linears are INT8 ConvRot. The global-attention layers keep Google's layout (see above).
  • Loading: needs the load-time patch shipped with our prompter node (see above).

gemma-4-31B-it_w4a8.safetensors

  • Role: W4A8 build of the official Gemma 4 31B-it.
  • Source: our conversion of gemma-4-31B-it_bf16.safetensors above (its name and sha256 are recorded in the file's __metadata__). The base model is by Google DeepMind.
  • License: Apache-2.0, as stated by Google (Gemma 4 license).
  • Changes: W4A8 quantization (recipe below); the other tensors are passed through in their source dtype.
  • Recipe: Comfy-Org's quant_int8_convrot.py from comfy-model-tools at commit 1846ff1a with --w4a8, and its own layer selection: a selected block linear whose input dimension is a multiple of 256 and that has at least 64 rows becomes asym_w4a8_int8 (4-bit weights with a codebook and FP8 scales per group of 16, ConvRot 256; activations quantized to INT8 at run time), the other selected linears and the token embedding become INT8 ConvRot, everything else is passed through.
  • Quantized layers: 600 (.comfy_quant entries in the file).
  • Quantizer check: the quantizer's own reconstruction check, dequantized and compared with the BF16 source: its 410 asym_w4a8_int8 linears relative error mean 7.32 % / max 7.35 % (the script computes no cosine for them); its 189 INT8 ConvRot linears cosine min 0.99995, relative error mean 0.82 % / max 1.03 %.
  • Layout (as published): the 410 language-model linears are asym_w4a8_int8 (group size 16, ConvRot 256); embed_tokens and the 189 vision-tower linears are INT8 ConvRot. The global-attention layers keep Google's layout (see above).
  • Requirements: comfy-kitchen's W4A8 kernels, which need a GPU with compute capability 8.0 or newer.
  • Loading: needs the load-time patch shipped with our prompter node (see above).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for beycanai/Gemma-LM

Finetuned
(196)
this model