AREX-2 27B · INT4 W4A16
Symmetric group-128 weight-only quantization for vLLM.
Base model · Paper · Project · vLLM
This is a numerical quantization of BAAI/AREX-2, not a fine-tune. It was exported from official revision
d4e3502f92d9e889031c2d04e387ab2eb3520268. Model, architecture, training-data, and upstream evaluation credit belongs to BAAI and the upstream contributors. Results from the upstream BF16 model do not measure this INT4 checkpoint.
Model summary
AREX-2 is a 27B-parameter, Qwen3.8-compatible multimodal model for long-horizon agent tasks. The upstream checkpoint declares a native context length of 262,144 tokens. Refer to the upstream model card for architecture, capabilities, evaluation protocols, and usage guidance.
Dual RTX 3090 deployment
Validation scope: the checkpoint loaded and passed a short text-chat smoke test in the pinned vLLM 0.29.0 deployment image on 2×RTX 3090, and completed the synthetic serving matrix below. No task-quality or multimodal evaluation was run; the native 262,144-token limit was not tested.
| Runtime property | Tested profile |
|---|---|
| GPUs | 2×RTX 3090 (Ampere sm_86) |
| Tensor parallelism | 2 |
| Activation dtype | BF16 |
| KV cache | FP8 E4M3 |
| CPU KV offload | 28 GiB, using the deployment's vLLM backport |
| Maximum model length | 262,144 configured; tested up to 128,003 prompt tokens + 1,024 generated tokens |
| Maximum active sequences | 1 |
| Optional speculation | External DFlash2 W4A16 drafter; not included in this repository |
Memory and capacity
| Metric | Result |
|---|---|
| Indexed tensor payload | 18.60 GB / 17.33 GiB |
| vLLM GPU memory after short probes | 16,784 MiB used and 7,394 MiB free per GPU |
| Long-context capacity | 128K prompt requests benchmarked; the configured 262,144-token limit was not tested |
GPU memory is the observed full serving-process allocation, not model weights alone. The longest benchmark request had 128,003 prompt tokens and 1,024 generated tokens; this does not establish capacity at the configured 262,144-token limit.
Measured serving performance
The matrix used synthetic, salted prompts at approximately 1K, 8K, 32K, 64K, and 128K tokens, with exact 512- or 1,024-token greedy completions. There were two cold-prefix runs per cell; TTFT is time to first streamed token, and prompt usage comes from the server. The table reports medians; the decode column also shows the two-run range. DFlash2 remained enabled (7 draft tokens, 15-token verify ceiling, lookup and chains on), with the production KV-cache and CPU-offload profile.
| Prompt tokens | New tokens | Median TTFT (s) | Prefill (tok/s) | Decode median (tok/s; min–max) |
|---|---|---|---|---|
| 1,027 | 512 | 0.69 | 1479 | 225.6 (191.3–260.0) |
| 1,028 | 1,024 | 0.71 | 1450 | 149.0 (88.0–209.9) |
| 8,194 | 512 | 5.73 | 1430 | 109.8 (75.9–143.7) |
| 8,196 | 1,024 | 5.79 | 1416 | 266.2 (253.9–278.4) |
| 32,002 | 512 | 25.17 | 1273 | 66.4 (65.2–67.7) |
| 32,002 | 1,024 | 28.71 | 1115 | 125.0 (63.3–186.8) |
| 64,002 | 512 | 74.38 | 861 | 78.5 (56.7–100.3) |
| 64,002 | 1,024 | 77.99 | 821 | 89.8 (57.6–122.0) |
| 128,002 | 512 | 180.24 | 710 | 42.6 (39.3–45.9) |
| 128,003 | 1,024 | 183.37 | 698 | 84.0 (81.6–86.4) |
Quantization fidelity
This export uses data-free, symmetric round-to-nearest (RTN) quantization. The aggregate relative L2 reconstruction error over the quantized weights is 0.12160. This is a weight-space diagnostic only; it is not perplexity, logit/KLD agreement, or a task-quality score. No calibration dataset, behavioral evaluation, or quality comparison against the BF16 source has been run.
Checkpoint profile
| Property | Value |
|---|---|
| Base checkpoint | BAAI/AREX-2, revision d4e3502f92d9e889031c2d04e387ab2eb3520268 |
| Quantization | Symmetric INT4 W4A16, group size 128, RTN; no calibration |
| Quantized modules | 400 two-dimensional Linear weights |
| Runtime format | compressed-tensors / pack-quantized |
| Kernel dispatch | CompressedTensorsWNA16 → MarlinLinearKernel in the tested runtime |
| Scale storage | FP16 |
| Preserved source precision | Vision tower, token embeddings, output head, recurrent GDN gates, and other excluded/non-Linear tensors |
| Runtime | vLLM with compressed-tensors support; this is not a GGUF checkpoint |
The quantization ignore rules are inherited from the validated W8A16 profile. The AREX-2 source checkpoint contains no MTP tensors; any speculative-decoding drafter must be supplied separately.
Quantization design
| Component | Treatment |
|---|---|
| 400 eligible two-dimensional Linear weights | Symmetric INT4, group size 128, round-to-nearest |
| Vision tower, embeddings, output head, recurrent GDN gates | Retained at source precision |
| Other non-Linear and excluded tensors | Retained at source precision |
The export uses the same exact eligible-module and ignore profile as the local W8A16 reference. It does not fine-tune or calibrate the model.
Why W4A16
W4A16 stores the selected Linear weights in 4-bit integers with group-wise scales while keeping activations at 16-bit precision. The smaller tensor payload does not guarantee higher inference speed or equivalent quality. Benchmark on the intended hardware and workload before choosing this variant for production.
Operational notes
| Question | Guidance |
|---|---|
| Which runtime is validated? | The pinned vLLM 0.29.0 deployment image used for the profile above. Other versions and kernels need separate validation. |
| Can llama.cpp load this checkpoint? | No. It uses the vLLM compressed-tensors format, not GGUF. |
| Is vision validated? | The source-precision vision weights are retained, but the smoke test was text-only; multimodal quality and memory use are unmeasured. |
| Is 262K serving validated? | No. The source model's context setting is retained, but long-context capacity and quality were not tested. |
| Is DFlash2 included? | No. The deployment smoke used an external drafter; this model repository contains only AREX-2 checkpoint assets. |
Files and provenance
The repository contains the sharded SafeTensors checkpoint, model.safetensors.index.json, model/tokenizer configuration, and the upstream Apache-2.0 LICENSE. Quantized tensors are represented by packed weights, per-group scales, and original weight shapes in the index.
Acknowledgements and license
This repository repackages numerical weights derived from BAAI/AREX-2. It does not claim authorship of the base model, architecture, training data, or upstream evaluations. The upstream Apache-2.0 license is included; retain the source attribution and follow the upstream terms.
- Downloads last month
- 399