Instructions to use hancheolp/nemo35_ternary with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use hancheolp/nemo35_ternary with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("hancheolp/nemo35_ternary") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use hancheolp/nemo35_ternary with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "hancheolp/nemo35_ternary" --prompt "Once upon a time"
- Atomic Chat
Nemotron 3.5 Lightning ternary (MLX 2-bit)
Ternary (1.58-bit) Nemotron 3.5 Lightning, post-training quantization only -- no distillation behind it. Mamba-2 plus mixture-of-experts, 128 experts per layer. Quantized: mixer in_proj/out_proj, q/k/v/o_proj, every expert up_proj and down_proj, shared experts. Full precision: A_log, D, dt_bias, conv1d, all norms, the MoE router, embeddings, lm_head, the MTP head. Expert down_proj has 1856 input channels, which 128 does not divide, so those modules use group_size 64.
Format
MLX native affine quantization, no custom kernel and no runtime shim:
bits 2
group_size 128
mode affine
levels {0, 1, 2} level 3 is unused
bias == -scale so dequantisation is scale * (q - 1) = {-a, 0, +a}
There is no rotation anywhere in this model, so there is no signs tensor and
no Hadamard transform to apply at load time.
Load
from mlx_lm import load
model, tokenizer = load("<repo>")
The container is stock MLX, but the architecture still has to be implemented in your mlx-lm / mlx-vlm build for the full model to load.
- Downloads last month
- -
2-bit