Nemotron 3.5 Lightning ternary (MLX 2-bit)

Ternary (1.58-bit) Nemotron 3.5 Lightning, post-training quantization only -- no distillation behind it. Mamba-2 plus mixture-of-experts, 128 experts per layer. Quantized: mixer in_proj/out_proj, q/k/v/o_proj, every expert up_proj and down_proj, shared experts. Full precision: A_log, D, dt_bias, conv1d, all norms, the MoE router, embeddings, lm_head, the MTP head. Expert down_proj has 1856 input channels, which 128 does not divide, so those modules use group_size 64.

Format

MLX native affine quantization, no custom kernel and no runtime shim:

bits        2
group_size  128
mode        affine
levels      {0, 1, 2}        level 3 is unused
bias        == -scale        so dequantisation is scale * (q - 1) = {-a, 0, +a}

There is no rotation anywhere in this model, so there is no signs tensor and no Hadamard transform to apply at load time.

Load

from mlx_lm import load
model, tokenizer = load("<repo>")

The container is stock MLX, but the architecture still has to be implemented in your mlx-lm / mlx-vlm build for the full model to load.

Downloads last month
-
Safetensors
Model size
33B params
Tensor type
F32
路
U32
路
BF16
路
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support