File size: 4,485 Bytes
a382da8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0feb39c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a382da8
 
 
 
0feb39c
a382da8
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
---
license: apache-2.0
base_model: nex-agi/Nex-N2.5-mini
tags:
- compressed-tensors
- int4
- int8
- qwen3_5_moe
- agentic
- vllm
- rtx-3090
- dgx-spark
pipeline_tag: image-text-to-text
---

# Nex-N2.5-mini Int4-Int8Mix

Data-free weight-only quantization of [nex-agi/Nex-N2.5-mini](https://huggingface.co/nex-agi/Nex-N2.5-mini) (a Qwen3.5-35B-A3B-family agentic model with a vision tower, 262K context) in the compressed-tensors `pack-quantized` format that stock vLLM loads directly. 66 GB of BF16 becomes 20 GB, which fits a pair of RTX 3090s at TP2 with room for a real KV pool, or one DGX Spark.

## Recipe

Same layer policy as our GLM-5.3 Int4-Int8Mix quant, adapted to this architecture:

| Tensors | Precision |
|---|---|
| 256 routed experts per layer, all 40 layers (`mlp.experts.N.gate_proj/up_proj/down_proj`) | INT4, symmetric, group 128 |
| Attention `q/k/v/o_proj`, Gated DeltaNet `in_proj_qkv`, `in_proj_z`, `out_proj`, shared expert `gate/up/down_proj` (layers 1 to 39) | INT8, symmetric, group 128 |
| Layer 0 side layers, routers (`mlp.gate`), `shared_expert_gate`, DeltaNet `in_proj_a`, `in_proj_b`, `conv1d`, `A_log`, `dt_bias`, all norms, embeddings, `lm_head`, the whole vision tower | BF16, untouched |

Round-to-nearest per group with min-max scales (compressed-tensors `QuantizationArgs`), no calibration data. The source checkpoint stores the experts fused per layer (`experts.gate_up_proj` [256, 1024, 2048], `experts.down_proj` [256, 2048, 512]); the quantizer splits them into per-expert 2-D weights and writes the standard per-expert names, which is what vLLM's Qwen3.5-MoE loader reads. Self-test round-trip error: INT4 experts about 11 to 14% mean relative per weight, INT8 layers about 0.8%.

Stats: 30,720 expert matrices INT4, 244 side-layer matrices INT8, 702 tensors passed through. 93,594 tensors, 16 shards, 19.6 GiB.

## Serving

Tested with vLLM v0.27.1 on 2x RTX 3090 (NVLink pair), Ampere, no FP8 compute:

```
vllm serve /models/Nex-N2.5-mini-Int4-Int8Mix \
  --served-model-name nex-n2.5-mini --port 8015 \
  --quantization compressed-tensors --dtype float16 \
  --tensor-parallel-size 2 --enable-expert-parallel \
  --max-model-len 262144 --gpu-memory-utilization 0.90 \
  --max-num-seqs 4 --max-num-batched-tokens 8192 \
  --kv-cache-dtype fp8_e5m2 --trust-remote-code --enable-chunked-prefill \
  --limit-mm-per-prompt '{"image":2,"audio":0}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --default-chat-template-kwargs '{"enable_thinking": false}'
```

Measured on that box (2026-09-08, single stream, temperature 0, thinking off, no draft):

| | |
|---|---|
| Weights per card | 9.98 GiB (Marlin WNA16 kernels) |
| KV pool | **1,867,521 tokens** at 262,144 context, 7.12x concurrency (fp8 e5m2 KV; 10 attention layers, 30 Gated DeltaNet layers) |
| Count to 100 | 185 tok/s, correct 1..100 |
| Prose (450 tokens) | 187 tok/s |
| Code (450 tokens) | 186 tok/s |
| Load | 27 s weights, 67 s torch.compile, serving in about 4 minutes |

This checkpoint has no MTP head, so no speculative decoding. Recipe, launcher, logs and the boot ladder: https://github.com/tonyd2wild/Nex-N2.5-mini-Int4-Int8Mix-2x-RTX-3090

### Config note (fixed 2026-09-08)

The `quantization_config` targets and ignore patterns are written as `re:.*layers[.]...` on purpose. vLLM matches quantization targets against its own module names (`language_model.model.layers.N...`), not the checkpoint names (`model.language_model.layers.N...`). The first upload pinned the patterns to the checkpoint prefix; vLLM then matched nothing, built the experts unquantized in fp16 and ran out of memory before loading a weight. The prefix-agnostic patterns match both namings, so the checkpoint loads in vLLM and in transformers.

## Caveats

- Data-free INT4 is a size play, not an accuracy claim; run your own evals for your workload.
- Quantized with our shard-streaming quantizer (`quant/nex_quant_stream.py` in the recipe repo, adapted from the GLM-5.3 Int4-Int8Mix recipe at https://github.com/tonyd2wild/GLM-5.3-Int4-Int8Mix-TP4-4x-DGX-Spark).

## Credits

- [nex-agi](https://huggingface.co/nex-agi) for Nex-N2.5-mini (Apache 2.0).
- Qwen team for the Qwen3.5 MoE architecture.
- [QuantTrio](https://huggingface.co/QuantTrio) for the Int4-Int8Mix layer policy we follow.
- vLLM and compressed-tensors (Neural Magic / Red Hat) for the format and kernels.

Quantized by Tony DeAngelo (Tech2Wild / 2Wild) with Kai, 2026-09-08.