AREX-2 EXL3 — 4.00 bpw
A 4.00 bpw ExLlamaV3 quantization of BAAI/AREX-2, a 27B dense multimodal agent model built on Qwen3.8-27B.
Source model: https://huggingface.co/BAAI/AREX-2
Paper: https://arxiv.org/abs/2609.38288
AREX-2 project: https://github.com/VectorSpaceLab/AREX-2
Why AREX-2
AREX-2 is designed around long-horizon iterative work rather than single-shot chat. Its training emphasizes cycles of proposing a solution, observing feedback, reflecting on failures, and revising the approach.
The source model is intended for:
- long-horizon agent workflows
- machine-learning engineering
- algorithmic coding
- tool-augmented deep research
- iterative self-improvement from scores, logs, errors, and other verifiable feedback
BAAI reports a 27B dense Qwen3.8-compatible multimodal architecture with a 262,144-token context window.
The benchmark numbers reported on the source model card belong to the original model. This EXL3 conversion has not been re-run across those full evaluation suites, so no claim is made that every BF16 benchmark score is preserved exactly at 4.00 bpw.
Quantization
| Item | Value |
|---|---|
| Source | BAAI/AREX-2 |
| Source revision | d4e3502f92d9e889031c2d04e387ab2eb3520268 |
| Architecture | Qwen3_5ForConditionalGeneration |
| Parameters | ~27.36B |
| Quantization | EXL3 |
| Target bitrate | 4.00 bpw |
| Head bits | 16 |
| Vision bits | 6 |
| ExLlamaV3 | 1.5.2+cu128.torch2.10.0 |
| Artifact size | ~17 GiB |
| License | Apache-2.0 |
The conversion was produced with the ExLlamaV3 conversion pipeline and validated after quantization rather than being published from conversion success alone.
Validation
This repository was tested locally on an RTX 4090 24 GB.
Direct ExLlamaV3 probe
| Check | Result |
|---|---|
| Model load | ✅ |
| Generation | ✅ |
| Load time | 2.416 s |
| Prefill | 9 tokens in 0.3327 s |
| Decode | 25 tokens in 0.6955 s |
| Decode throughput | 35.94 tok/s |
| Torch | 2.10.0+cu128 |
The throughput figure is a small local smoke benchmark, not a standardized cross-system performance result.
TabbyAPI smoke matrix
Validated with an 8,192-token runtime context and batch size 2:
| Capability | Result |
|---|---|
| OpenAI-compatible model endpoint | ✅ |
| Standard chat completion | ✅ |
| Streaming completion | ✅ |
| Tool calling | ✅ |
| Vision input | ✅ |
| Two concurrent requests | ✅ |
Observed standard-chat throughput during the TabbyAPI smoke test was about 49.4 tok/s on the same RTX 4090. This was a very short request and should not be treated as a sustained benchmark.
Tool calling was validated with a structured function call, and vision was validated with an image question rather than by checking only whether the multimodal modules loaded.
Reasoning behavior
AREX-2's chat template is reasoning-aware. For ordinary chat, enabling thinking is the safer default.
A useful low-overhead setting is:
{
"chat_template_kwargs": {
"enable_thinking": true,
"reasoning_effort": "low"
}
}
During validation, forcing enable_thinking=false could cause a trivial exact-answer prompt to terminate immediately with no visible content. With thinking enabled, normal chat completed correctly.
That behavior is worth knowing if a frontend or serving preset disables reasoning globally.
Running with TabbyAPI
Point TabbyAPI at the downloaded model directory and use the ExLlamaV3 backend.
The validation configuration used:
model:
backend: exllamav3
max_seq_len: 8192
cache_size: 8192
cache_mode: FP16
max_batch_size: 2
vision: true
reasoning: true
The source model advertises a much larger context window, but this quant was only smoke-tested at 8,192 tokens. Long-context quality and memory scaling at 32K/64K/262K were not validated for this release.
What was and was not tested
Tested:
- EXL3 model load
- text generation
- OpenAI-compatible chat serving
- streaming
- tool-call parsing
- multimodal/vision path
- two-request concurrency
Not yet tested:
- full source benchmark reproduction
- BF16 vs EXL3 quality delta
- 32K+ context retention
- long-horizon AREX task loops under quantization
- sustained throughput under larger batch sizes
If you care about agentic quality rather than just whether the model loads, the most useful next test is a BF16-vs-4bpw comparison on multi-round reflective tasks, because that is the behavior AREX-2 is specifically designed around.
Provenance
This quantization was generated from:
BAAI/AREX-2@d4e3502f92d9e889031c2d04e387ab2eb3520268
The published weights were validated locally before upload. The source model remains subject to its original model card, license, intended-use guidance, and upstream notices.
Source citation
@article{2026arex2,
title = {AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks},
author = {Qian, Hongjin and Li, Chaofan and Luo, Kun and Wei, Wenqing and Chen, Jianlyu and Lu, Shuqi and Hu, Yuyang and Xiao, Hongwang and Wang, Hui and Li, Chaozhuo and Ye, Qiwei and Dou, Zhicheng and Lian, Defu and Liu, Zheng},
journal = {arXiv preprint arXiv:2609.38288},
year = {2026},
url = {https://arxiv.org/abs/2609.38288}
}
- Downloads last month
- 35