nina-arm commited on
Commit
3712556
·
verified ·
1 Parent(s): b840f3a

Sync model repo (text/metadata)

Browse files
Files changed (2) hide show
  1. README.md +18 -18
  2. metadata.yaml +0 -20
README.md CHANGED
@@ -1,10 +1,11 @@
1
  ---
2
  library_name: executorch
3
- display_name: Gemma-4-E2B 8da4w — ExecuTorch
4
  license: apache-2.0
5
  tags:
6
  - text-generation
7
  - 8da4w
 
8
  - quantized
9
  - xnnpack
10
  - kleidiai
@@ -12,13 +13,13 @@ tags:
12
  - executorch
13
  - edge-ai
14
  pipeline_tag: text-generation
15
- base_model: google/gemma-4-E2B
16
  base_model_relation: quantized
17
  ---
18
 
19
- # Gemma-4-E2B 8da4w (ExecuTorch + XNNPACK)
20
- This is an **8da4w-quantized** (8-bit dynamic per-token activations + 4-bit per-channel grouped weights) version of [google/gemma-4-E2B](https://huggingface.co/google/gemma-4-E2B), optimized for edge deployment on ARM devices using [ExecuTorch](https://github.com/pytorch/executorch) with the **XNNPACK + KleidiAI** backend.
21
- The model was quantized by ExecuTorch’s Gemma 4 export script without SDPA, with a custom static KV-cache, and exported to the `.pte` format for efficient on-device inference on ARM Cortex-A processors (modern Android phones, AWS Graviton, embedded ARM).
22
 
23
  ## Key Highlights
24
 
@@ -30,8 +31,8 @@ Compared to the FP32 baseline:
30
 
31
  ### Model Description
32
 
33
- Gemma 4 E2B-IT is an instruction-tuned multimodal generative model from Google DeepMind. This text-only artifact is the same model after 8da4w post-training quantization and ExecuTorch/XNNPACK export, no fine-tuning was performed.
34
- Linear weights use groupwise INT4 quantization, activations are dynamically quantized to INT8, and embeddings use INT8.
35
 
36
  - **Developed by:** Google DeepMind (base model); Arm Model Optimization pipeline (quantization and export)
37
  - **Model type:** Text generation (autoregressive language model)
@@ -40,7 +41,7 @@ Linear weights use groupwise INT4 quantization, activations are dynamically quan
40
 
41
  ### Model Sources
42
 
43
- - **Repository:** https://huggingface.co/google/gemma-4-E2B
44
  - **Paper:** [Gemma 4 Technical Report](https://arxiv.org/abs/2607.02770)
45
 
46
  ## How to Get Started with the Model
@@ -61,35 +62,34 @@ Linear weights use groupwise INT4 quantization, activations are dynamically quan
61
 
62
  ### Objective
63
 
64
- Open-ended text generation and multiple-choice scoring with the Gemma 4 E2B architecture, deployed on ARM CPUs using INT4 linear weights and INT8 activations.
65
 
66
  ### Quantization
67
 
68
  - **Activations**: 8-bit dynamic, per-token quantization
69
- - **Linear weights**: 4-bit, per-channel grouped quantization with group size 32
70
- - **Token embedding**: 8-bit, per-channel weight quantization
71
  - **Backend**: XNNPACK with KleidiAI INT4 micro-kernels
72
- - **Higher-precision layers**: None. Every linear layer and the token embedding are quantized. Attention internals remain in higher precision within the custom fused SDPA path.
73
- - **Calibration**: 200 prompts randomly sampled from the MMLU test split
74
 
75
  ### Export Pipeline
76
 
77
  1. Load the pretrained Gemma 4 E2B-IT weights from Hugging Face safetensors and convert them into ExecuTorch’s custom Gemma 4 text-decoder architecture.
78
  2. Configure a static KV cache with a maximum sequence length of 1,024 tokens and enable dynamically sized input sequences.
79
- 3. Replace standard attention with the fused llama::custom_sdpa tiled-attention operator.
80
  4. Apply TorchAO 8da4w post-training quantization:
81
  - dynamic INT8 activations
82
  - grouped INT4 linear weights
83
  - group size 128
84
  - incompatible layers retained at higher precision
85
- 5. Quantize embedding tables to INT8. No calibration dataset or observer-calibration pass is used.
86
  6. Capture the quantized decoder with `torch.export` and decompose it into an ExecuTorch-compatible graph.
87
  7. Lower supported quantized and floating-point operations to the XNNPACK delegate and apply static memory planning.
88
  8. Serialize the graph and weights as an ExecuTorch `.pte` artifact compatible with the generic LLM runner.
89
 
90
  ## Known Limitations
91
 
92
- - This is a text-only export. The original Gemma 4 E2B model’s image and audio encoders are not included.
93
- - 8da4w post-training quantization can reduce accuracy relative to the original model. Linear weights are INT4, dynamic activations are INT8, and embeddings are INT8.
94
  - Initial inference may incur one-time weight preparation or repacking overhead. Subsequent inference can be faster while the model remains loaded.
95
-
 
1
  ---
2
  library_name: executorch
3
+ display_name: Gemma-4-E2B-IT 8da4w+emb4 — ExecuTorch
4
  license: apache-2.0
5
  tags:
6
  - text-generation
7
  - 8da4w
8
+ - emb4
9
  - quantized
10
  - xnnpack
11
  - kleidiai
 
13
  - executorch
14
  - edge-ai
15
  pipeline_tag: text-generation
16
+ base_model: google/gemma-4-E2B-it
17
  base_model_relation: quantized
18
  ---
19
 
20
+ # Gemma-4-E2B-IT 8da4w+emb4 (ExecuTorch + XNNPACK)
21
+ This is an **8da4w+emb4-quantized** (8-bit dynamic per-token activations + 4-bit per-channel grouped linear weights + packed 4-bit embedding weights) version of [google/gemma-4-E2B-it](https://huggingface.co/google/gemma-4-E2B-it), optimized for edge deployment on ARM devices using [ExecuTorch](https://github.com/pytorch/executorch) with the **XNNPACK + KleidiAI** backend.
22
+ The model was quantized by ExecuTorch’s Gemma 4 export script with custom fused SDPA and a static KV cache, then exported to the `.pte` format for efficient on-device inference on ARM Cortex-A processors (modern Android phones, AWS Graviton, embedded ARM).
23
 
24
  ## Key Highlights
25
 
 
31
 
32
  ### Model Description
33
 
34
+ Gemma 4 E2B-IT is an instruction-tuned multimodal generative model from Google DeepMind. This text-only artifact is the same model after 8da4w+emb4 post-training quantization and ExecuTorch/XNNPACK export; no fine-tuning was performed.
35
+ Linear weights use groupwise INT4 quantization, activations are dynamically quantized to INT8, and embeddings use packed INT4 weights.
36
 
37
  - **Developed by:** Google DeepMind (base model); Arm Model Optimization pipeline (quantization and export)
38
  - **Model type:** Text generation (autoregressive language model)
 
41
 
42
  ### Model Sources
43
 
44
+ - **Repository:** https://huggingface.co/google/gemma-4-E2B-it
45
  - **Paper:** [Gemma 4 Technical Report](https://arxiv.org/abs/2607.02770)
46
 
47
  ## How to Get Started with the Model
 
62
 
63
  ### Objective
64
 
65
+ Open-ended text generation and multiple-choice scoring with the instruction-tuned Gemma 4 E2B architecture, deployed on ARM CPUs using INT4 linear weights and INT8 activations.
66
 
67
  ### Quantization
68
 
69
  - **Activations**: 8-bit dynamic, per-token quantization
70
+ - **Linear weights**: 4-bit grouped quantization with group size 128
71
+ - **Token embedding**: packed 4-bit weight quantization
72
  - **Backend**: XNNPACK with KleidiAI INT4 micro-kernels
73
+ - **Higher-precision operations**: Custom fused SDPA keeps attention internals in floating point. Linear layers with shapes incompatible with grouped INT4 quantization, if any, are retained at their original precision.
74
+ - **Calibration**: None. Dynamic activation quantization does not use a calibration dataset or observer-calibration pass.
75
 
76
  ### Export Pipeline
77
 
78
  1. Load the pretrained Gemma 4 E2B-IT weights from Hugging Face safetensors and convert them into ExecuTorch’s custom Gemma 4 text-decoder architecture.
79
  2. Configure a static KV cache with a maximum sequence length of 1,024 tokens and enable dynamically sized input sequences.
80
+ 3. Replace standard attention with the fused `llama::custom_sdpa` tiled-attention operator.
81
  4. Apply TorchAO 8da4w post-training quantization:
82
  - dynamic INT8 activations
83
  - grouped INT4 linear weights
84
  - group size 128
85
  - incompatible layers retained at higher precision
86
+ 5. Quantize embedding tables to packed INT4 weights. No calibration dataset or observer-calibration pass is used.
87
  6. Capture the quantized decoder with `torch.export` and decompose it into an ExecuTorch-compatible graph.
88
  7. Lower supported quantized and floating-point operations to the XNNPACK delegate and apply static memory planning.
89
  8. Serialize the graph and weights as an ExecuTorch `.pte` artifact compatible with the generic LLM runner.
90
 
91
  ## Known Limitations
92
 
93
+ - This is a text-only export. The original Gemma 4 E2B-IT model’s image and audio encoders are not included.
94
+ - 8da4w+emb4 post-training quantization can reduce accuracy relative to the original model. Linear and embedding weights are INT4, and dynamic activations are INT8.
95
  - Initial inference may incur one-time weight preparation or repacking overhead. Subsequent inference can be faster while the model remains loaded.
 
metadata.yaml DELETED
@@ -1,20 +0,0 @@
1
- schema_version: 1.0.0
2
- report_type: llm-generative
3
- task_type: text-generation
4
- title: Gemma-4-E2B 8da4w — ExecuTorch
5
- id: Arm/gemma-4-e2b-text-executorch
6
- base_model_id: google/gemma-4-E2B
7
- vendor: Google
8
- base_model_url: https://huggingface.co/google/gemma-4-E2B
9
- profile: Arm-Optimized
10
- weight_dtype: int4
11
- model_size_mb: 2563.959
12
- format: executorch
13
- filename: gemma4_e2b_text_8da4w_emb4.pte
14
- quantization:
15
- method: PTQ-dynamic
16
- variant: 8da4w
17
- weight_bits: 4
18
- activation_bits: 8
19
- mode: dynamic
20
- weight_granularity: per-group