Model Card for Gemma-3n-E4B-IT (Aria Quant Bundle, q4)

Model Details

Model Description

Gemma-3n-E4B-IT is a ~4-billion-parameter instruction-tuned multimodal language model developed by Google, part of the Gemma 3n family. Its text backbone features hybrid attention (28 of 35 layers sliding-window linear attention with activation sparsity + 7 Laurel low-rank full-attention layers), GeGLU activation, Grouped Query Attention (GQA), per-layer input projections, and 32K native context length. Pre-trained on diverse web-scale corpora and aligned via instruction tuning + RLHF. This distribution is provided by Aria Compute as an aria-quant-bundle — a quantized package using Hadamard pre-processing + per-channel 4-bit quantization. Optimized for CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the Aria Engine runtime. No GPU or cloud connection is required.

  • Developed by: Google
  • Quantized and distributed by: Aria Compute
  • Model type: Dense Transformer decoder-only (multimodal base: image/audio + text inputs, text outputs; this bundle ships the text backbone)
  • Language(s): English (primary), Chinese, and 30+ additional languages
  • License: Apache 2.0
  • Finetuned from model: google/gemma-3n-e4b-it

Model Sources

Uses

Direct Use

This quantized bundle is intended for on-device, offline text-generation tasks on resource-constrained hardware, including:

  • On-device chat and conversational assistants
  • Real-time text completion and basic code snippet generation
  • Instruction-following tasks for mobile and IoT applications
  • Lightweight text embeddings for on-device retrieval and classification
  • Short-form summarization of notifications, messages, and local content
  • Local document analysis up to 32K context (chunked)

All inference runs locally on CPU. No data is sent to external servers.

Target Devices

Platform Runtime Memory Feasibility
High-end smartphone (8 GB) ~1.0 GB ✅ Recommended
Mid-range smartphone (4–6 GB) ~1.0 GB
Budget phone (2–3 GB) ~1.0 GB
Raspberry Pi 5 / SBC (4–8 GB) ~1.0 GB
IoT gateway (1–2 GB) ~1.0 GB ⚠️ Tight fit
Wearable (1 GB) ~1.0 GB

Memory breakdown (q4, at 4K context): ~0.98 GB quantized model weights (mmap) + ~80 MB KV cache + ~30 MB runtime overhead + ~45 MB per-channel metadata overhead ≈ ~1.1 GB.

Note: KV cache is compact thanks to hybrid attention — 28 of 35 layers use sliding-window attention (window 512, KV bounded by the window), so only the 7 Laurel full-attention layers scale KV with context. Combined with GQA (2 KV heads), this keeps 32K context practical on ~2 GB-class devices.

Out-of-Scope Use

  • Long-form creative writing (>2K tokens per generation)
  • Mathematical theorem proving or formal verification
  • Full program/application synthesis
  • Multimodal input (image/audio encoding pipeline is pending audit for this quantized bundle — text-only in this release)
  • Real-time audio/speech processing (use Aria speech models)
  • Safety-critical decision systems without human oversight
  • Deployment in production when batch inference or GPU acceleration is required (this bundle targets CPU-only, single-prompt inference)
  • Tasks requiring factual precision beyond the model's ~4B parameter capacity

How to Get Started with the Model

Download from Aria Compute

Authenticated dashboard users can download the bundle via: https://ariacompute.com/dashboard/models

Quantization Recipe

This bundle uses a per-channel quantization recipe, one of several precision options in the Aria Compute lineup:

Component Quantization Strategy Details
Attention Q/K/V/O weights 4-bit Uniform per-channel codebooks, Hadamard pre-processing
FFN gate/up/down weights 4-bit Uniform per-channel codebooks, Hadamard pre-processing
RMSNorm weights FP16 Preserved at full precision
Embedding table FP16 Preserved at full precision
  • Bundle size: ~1.0 GB (BF16 text backbone: ~7.8 GB)
  • Generation quality: Awaiting gen_quant_eval audit. Uniform 4-bit per-channel quantization is the small-footprint baseline in the Aria lineup — smallest bundle with full per-channel codebook compression. Per-channel codebooks preserve per-output-channel distribution characteristics. Formal quality benchmarks against FP16 and other Aria quant recipes are pending
  • Calibration-free: Hadamard pre-processing + per-channel quantization, no task-specific calibration data required
  • Other precision options: Also available for Gemma-3n-E4B-IT: gemma-3n-e4b-it_q8_channel (per-channel 8-bit, near-lossless) and gemma-3n-e4b-it_q326_channel (mixed precision, attn 4-bit + FFN ~3-bit, recommended quality-size trade-off)

Model Architecture

Gemma-3n-E4B-IT's text backbone employs a dense Transformer decoder with GeGLU activation, hybrid attention (sliding-window linear attention + Laurel low-rank full attention), and per-layer input projections (256-dim layer input → 2,048 hidden):

Parameter Gemma 3n E4B
Layers 35
Hidden size 2,048
Layer input size 256
FFN intermediate size 16,384
Attention heads (Query) 8
Attention heads (KV) 2 (GQA)
Head dimension 256
Full-attention (Laurel) layers 7
Sliding-window layers 28
Sliding window 512
Laurel rank 64
Activation GeGLU
Position encoding RoPE (θ = 1,000,000)
Normalization RMSNorm (pre-norm)
Vocabulary size 262,400
Max context length 32,768

Design highlights (Gemma 3n family):

  • Hybrid attention: 28/35 layers use sliding-window attention (window 512) with activation sparsity (first 10 layers at 95% sparsity); only 7/35 layers use full softmax attention with Laurel (low-rank attention, rank 64) — dense-attention KV cost is confined to a few layers
  • GQA (Grouped Query Attention): 2 KV heads serving 8 query heads — KV Cache memory is 4× smaller than full attention
  • Per-layer input projections: a 256-dim layer-input embedding is projected to the 2,048-dim hidden state at each layer — compact embedding table with shared layer-input processing
  • GeGLU activation: GELU with tanh approximation gating, efficient for on-device inference
  • RoPE position encoding: 1M base frequency (global) with 10K local base frequency for sliding-window layers, supporting 32K context
  • RMSNorm pre-normalization: Lightweight normalization before each sub-layer
  • Final logit softcapping: Output logits capped at ±30.0 for training stability

Bias, Risks, and Limitations

Limitations

  • Reasoning depth: Multi-step logical reasoning is limited for a 2B-class model. Verify outputs in high-stakes scenarios; consider larger Gemma 3n variants for reasoning tasks.
  • Mathematics: Simple arithmetic may be attempted but is unreliable. Advanced quantitative reasoning is out of scope. Use larger models for mathematical tasks.
  • Code generation: Capable of single-line completions and basic snippets; unreliable for multi-line code or structured programs.
  • Factual knowledge: Limited world knowledge due to ~4B parameter scale. Always verify factual claims against authoritative sources. This model is best suited for instruction-following and lightweight text processing rather than encyclopedic knowledge retrieval.
  • Instruction following: Handles simple single-constraint instructions. Complex multi-constraint prompts may cause degradation, especially at longer contexts.
  • Quantization drift: 4-bit per-channel quantization may exhibit more generation drift versus FP16 than higher-precision recipes, especially on ambiguous or open-ended prompts. For better fidelity, use gemma-3n-e4b-it_q326_channel (mixed precision, attn 4-bit + FFN ~3-bit) or gemma-3n-e4b-it_q8_channel (per-channel 8-bit).

Bias and Risks

  • Bias: As with all large language models trained on web-scale data, Gemma 3n may reflect societal biases present in its training corpus. Evaluate outputs before deployment in sensitive domains (hiring, healthcare, law).
  • Toxicity: The instruction-tuned model has been safety-aligned with RLHF. However, no safety filter is exhaustive. Consider an additional output classifier in high-risk environments.
  • Hallucination: May generate plausible-sounding but factually incorrect information. Implement output verification for critical applications. Hallucination risk is elevated for smaller models due to limited memorization capacity.
  • Dual-use risk: Text-generation capabilities could be misused for spam, disinformation, or impersonation. Deploy responsibly and in accordance with the Apache 2.0 license terms.

Recommendations

Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend:

  • Adding a lightweight output safety classifier for user-facing deployments
  • Verifying factual claims with external knowledge bases
  • Not using the model for high-stakes decisions without human review
  • Considering larger Gemma 3n variants (E4B) for tasks requiring stronger reasoning or factual recall
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train ariacompute/gemma-3n-e4b-it_q4

Paper for ariacompute/gemma-3n-e4b-it_q4

Evaluation results