Sync model repo (text/metadata)
Browse files- README.md +18 -18
- metadata.yaml +0 -20
README.md
CHANGED
|
@@ -1,10 +1,11 @@
|
|
| 1 |
---
|
| 2 |
library_name: executorch
|
| 3 |
-
display_name: Gemma-4-E2B 8da4w — ExecuTorch
|
| 4 |
license: apache-2.0
|
| 5 |
tags:
|
| 6 |
- text-generation
|
| 7 |
- 8da4w
|
|
|
|
| 8 |
- quantized
|
| 9 |
- xnnpack
|
| 10 |
- kleidiai
|
|
@@ -12,13 +13,13 @@ tags:
|
|
| 12 |
- executorch
|
| 13 |
- edge-ai
|
| 14 |
pipeline_tag: text-generation
|
| 15 |
-
base_model: google/gemma-4-E2B
|
| 16 |
base_model_relation: quantized
|
| 17 |
---
|
| 18 |
|
| 19 |
-
# Gemma-4-E2B 8da4w (ExecuTorch + XNNPACK)
|
| 20 |
-
This is an **8da4w-quantized** (8-bit dynamic per-token activations + 4-bit per-channel grouped weights) version of [google/gemma-4-E2B](https://huggingface.co/google/gemma-4-E2B), optimized for edge deployment on ARM devices using [ExecuTorch](https://github.com/pytorch/executorch) with the **XNNPACK + KleidiAI** backend.
|
| 21 |
-
The model was quantized by ExecuTorch’s Gemma 4 export script
|
| 22 |
|
| 23 |
## Key Highlights
|
| 24 |
|
|
@@ -30,8 +31,8 @@ Compared to the FP32 baseline:
|
|
| 30 |
|
| 31 |
### Model Description
|
| 32 |
|
| 33 |
-
Gemma 4 E2B-IT is an instruction-tuned multimodal generative model from Google DeepMind. This text-only artifact is the same model after 8da4w post-training quantization and ExecuTorch/XNNPACK export
|
| 34 |
-
Linear weights use groupwise INT4 quantization, activations are dynamically quantized to INT8, and embeddings use
|
| 35 |
|
| 36 |
- **Developed by:** Google DeepMind (base model); Arm Model Optimization pipeline (quantization and export)
|
| 37 |
- **Model type:** Text generation (autoregressive language model)
|
|
@@ -40,7 +41,7 @@ Linear weights use groupwise INT4 quantization, activations are dynamically quan
|
|
| 40 |
|
| 41 |
### Model Sources
|
| 42 |
|
| 43 |
-
- **Repository:** https://huggingface.co/google/gemma-4-E2B
|
| 44 |
- **Paper:** [Gemma 4 Technical Report](https://arxiv.org/abs/2607.02770)
|
| 45 |
|
| 46 |
## How to Get Started with the Model
|
|
@@ -61,35 +62,34 @@ Linear weights use groupwise INT4 quantization, activations are dynamically quan
|
|
| 61 |
|
| 62 |
### Objective
|
| 63 |
|
| 64 |
-
Open-ended text generation and multiple-choice scoring with the Gemma 4 E2B architecture, deployed on ARM CPUs using INT4 linear weights and INT8 activations.
|
| 65 |
|
| 66 |
### Quantization
|
| 67 |
|
| 68 |
- **Activations**: 8-bit dynamic, per-token quantization
|
| 69 |
-
- **Linear weights**: 4-bit
|
| 70 |
-
- **Token embedding**:
|
| 71 |
- **Backend**: XNNPACK with KleidiAI INT4 micro-kernels
|
| 72 |
-
- **Higher-precision
|
| 73 |
-
- **Calibration**:
|
| 74 |
|
| 75 |
### Export Pipeline
|
| 76 |
|
| 77 |
1. Load the pretrained Gemma 4 E2B-IT weights from Hugging Face safetensors and convert them into ExecuTorch’s custom Gemma 4 text-decoder architecture.
|
| 78 |
2. Configure a static KV cache with a maximum sequence length of 1,024 tokens and enable dynamically sized input sequences.
|
| 79 |
-
3. Replace standard attention with the fused llama::custom_sdpa tiled-attention operator.
|
| 80 |
4. Apply TorchAO 8da4w post-training quantization:
|
| 81 |
- dynamic INT8 activations
|
| 82 |
- grouped INT4 linear weights
|
| 83 |
- group size 128
|
| 84 |
- incompatible layers retained at higher precision
|
| 85 |
-
5. Quantize embedding tables to
|
| 86 |
6. Capture the quantized decoder with `torch.export` and decompose it into an ExecuTorch-compatible graph.
|
| 87 |
7. Lower supported quantized and floating-point operations to the XNNPACK delegate and apply static memory planning.
|
| 88 |
8. Serialize the graph and weights as an ExecuTorch `.pte` artifact compatible with the generic LLM runner.
|
| 89 |
|
| 90 |
## Known Limitations
|
| 91 |
|
| 92 |
-
- This is a text-only export. The original Gemma 4 E2B model’s image and audio encoders are not included.
|
| 93 |
-
- 8da4w post-training quantization can reduce accuracy relative to the original model. Linear weights are INT4, dynamic activations are INT8
|
| 94 |
- Initial inference may incur one-time weight preparation or repacking overhead. Subsequent inference can be faster while the model remains loaded.
|
| 95 |
-
|
|
|
|
| 1 |
---
|
| 2 |
library_name: executorch
|
| 3 |
+
display_name: Gemma-4-E2B-IT 8da4w+emb4 — ExecuTorch
|
| 4 |
license: apache-2.0
|
| 5 |
tags:
|
| 6 |
- text-generation
|
| 7 |
- 8da4w
|
| 8 |
+
- emb4
|
| 9 |
- quantized
|
| 10 |
- xnnpack
|
| 11 |
- kleidiai
|
|
|
|
| 13 |
- executorch
|
| 14 |
- edge-ai
|
| 15 |
pipeline_tag: text-generation
|
| 16 |
+
base_model: google/gemma-4-E2B-it
|
| 17 |
base_model_relation: quantized
|
| 18 |
---
|
| 19 |
|
| 20 |
+
# Gemma-4-E2B-IT 8da4w+emb4 (ExecuTorch + XNNPACK)
|
| 21 |
+
This is an **8da4w+emb4-quantized** (8-bit dynamic per-token activations + 4-bit per-channel grouped linear weights + packed 4-bit embedding weights) version of [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it), optimized for edge deployment on ARM devices using [ExecuTorch](https://github.com/pytorch/executorch) with the **XNNPACK + KleidiAI** backend.
|
| 22 |
+
The model was quantized by ExecuTorch’s Gemma 4 export script with custom fused SDPA and a static KV cache, then exported to the `.pte` format for efficient on-device inference on ARM Cortex-A processors (modern Android phones, AWS Graviton, embedded ARM).
|
| 23 |
|
| 24 |
## Key Highlights
|
| 25 |
|
|
|
|
| 31 |
|
| 32 |
### Model Description
|
| 33 |
|
| 34 |
+
Gemma 4 E2B-IT is an instruction-tuned multimodal generative model from Google DeepMind. This text-only artifact is the same model after 8da4w+emb4 post-training quantization and ExecuTorch/XNNPACK export; no fine-tuning was performed.
|
| 35 |
+
Linear weights use groupwise INT4 quantization, activations are dynamically quantized to INT8, and embeddings use packed INT4 weights.
|
| 36 |
|
| 37 |
- **Developed by:** Google DeepMind (base model); Arm Model Optimization pipeline (quantization and export)
|
| 38 |
- **Model type:** Text generation (autoregressive language model)
|
|
|
|
| 41 |
|
| 42 |
### Model Sources
|
| 43 |
|
| 44 |
+
- **Repository:** https://huggingface.co/google/gemma-4-E2B-it
|
| 45 |
- **Paper:** [Gemma 4 Technical Report](https://arxiv.org/abs/2607.02770)
|
| 46 |
|
| 47 |
## How to Get Started with the Model
|
|
|
|
| 62 |
|
| 63 |
### Objective
|
| 64 |
|
| 65 |
+
Open-ended text generation and multiple-choice scoring with the instruction-tuned Gemma 4 E2B architecture, deployed on ARM CPUs using INT4 linear weights and INT8 activations.
|
| 66 |
|
| 67 |
### Quantization
|
| 68 |
|
| 69 |
- **Activations**: 8-bit dynamic, per-token quantization
|
| 70 |
+
- **Linear weights**: 4-bit grouped quantization with group size 128
|
| 71 |
+
- **Token embedding**: packed 4-bit weight quantization
|
| 72 |
- **Backend**: XNNPACK with KleidiAI INT4 micro-kernels
|
| 73 |
+
- **Higher-precision operations**: Custom fused SDPA keeps attention internals in floating point. Linear layers with shapes incompatible with grouped INT4 quantization, if any, are retained at their original precision.
|
| 74 |
+
- **Calibration**: None. Dynamic activation quantization does not use a calibration dataset or observer-calibration pass.
|
| 75 |
|
| 76 |
### Export Pipeline
|
| 77 |
|
| 78 |
1. Load the pretrained Gemma 4 E2B-IT weights from Hugging Face safetensors and convert them into ExecuTorch’s custom Gemma 4 text-decoder architecture.
|
| 79 |
2. Configure a static KV cache with a maximum sequence length of 1,024 tokens and enable dynamically sized input sequences.
|
| 80 |
+
3. Replace standard attention with the fused `llama::custom_sdpa` tiled-attention operator.
|
| 81 |
4. Apply TorchAO 8da4w post-training quantization:
|
| 82 |
- dynamic INT8 activations
|
| 83 |
- grouped INT4 linear weights
|
| 84 |
- group size 128
|
| 85 |
- incompatible layers retained at higher precision
|
| 86 |
+
5. Quantize embedding tables to packed INT4 weights. No calibration dataset or observer-calibration pass is used.
|
| 87 |
6. Capture the quantized decoder with `torch.export` and decompose it into an ExecuTorch-compatible graph.
|
| 88 |
7. Lower supported quantized and floating-point operations to the XNNPACK delegate and apply static memory planning.
|
| 89 |
8. Serialize the graph and weights as an ExecuTorch `.pte` artifact compatible with the generic LLM runner.
|
| 90 |
|
| 91 |
## Known Limitations
|
| 92 |
|
| 93 |
+
- This is a text-only export. The original Gemma 4 E2B-IT model’s image and audio encoders are not included.
|
| 94 |
+
- 8da4w+emb4 post-training quantization can reduce accuracy relative to the original model. Linear and embedding weights are INT4, and dynamic activations are INT8.
|
| 95 |
- Initial inference may incur one-time weight preparation or repacking overhead. Subsequent inference can be faster while the model remains loaded.
|
|
|
metadata.yaml
DELETED
|
@@ -1,20 +0,0 @@
|
|
| 1 |
-
schema_version: 1.0.0
|
| 2 |
-
report_type: llm-generative
|
| 3 |
-
task_type: text-generation
|
| 4 |
-
title: Gemma-4-E2B 8da4w — ExecuTorch
|
| 5 |
-
id: Arm/gemma-4-e2b-text-executorch
|
| 6 |
-
base_model_id: google/gemma-4-E2B
|
| 7 |
-
vendor: Google
|
| 8 |
-
base_model_url: https://huggingface.co/google/gemma-4-E2B
|
| 9 |
-
profile: Arm-Optimized
|
| 10 |
-
weight_dtype: int4
|
| 11 |
-
model_size_mb: 2563.959
|
| 12 |
-
format: executorch
|
| 13 |
-
filename: gemma4_e2b_text_8da4w_emb4.pte
|
| 14 |
-
quantization:
|
| 15 |
-
method: PTQ-dynamic
|
| 16 |
-
variant: 8da4w
|
| 17 |
-
weight_bits: 4
|
| 18 |
-
activation_bits: 8
|
| 19 |
-
mode: dynamic
|
| 20 |
-
weight_granularity: per-group
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|