Nemotron 3.5 Lightning 30B-A3B ternary (MLX 2-bit)

Ternary (1.58-bit) Nemotron 3.5 Lightning 30B-A3B. Same module layout as the mlx-community 4-bit release: the MoE router and every norm, A_log, D, dt_bias and conv1d stay full precision, and the experts ship pre-stacked as switch_mlp.fc1/fc2. The body weights are ternary. Embeddings are affine 4-bit and lm_head is affine 8-bit: the chat stop token's row carries a third of the median norm, which puts it in the worst 1% of rows at 4 bits.

Format

MLX native affine quantization, no custom kernel and no runtime shim:

bits        2
group_size  64
mode        affine
levels      {0, 1, 2}        level 3 is unused
bias        == -scale        so dequantisation is scale * (q - 1) = {-a, 0, +a}

There is no rotation anywhere in this model, so there is no signs tensor and no Hadamard transform to apply at load time.

Load

from mlx_lm import load
model, tokenizer = load("<repo>")

The container is stock MLX, but the architecture still has to be implemented in your mlx-lm / mlx-vlm build for the full model to load.

Downloads last month
-
Safetensors
Model size
33B params
Tensor type
F32
路
U32
路
BF16
路
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for hancheolp/nemo_final

Quantized
(109)
this model