Index-Translate-2B-MLX-8bit

MLX affine 8-bit quantization (group size 64) of IndexTeam/Index-Translate-2B. Weights only re-encoded — no retraining, no fine-tuning.

Text translation only: the upstream checkpoint's vision tower and MTP head are not included. Runs with mlx-lm on Apple Silicon; not a llama.cpp / GGUF model.

Usage

from mlx_lm import load, generate

model, tokenizer = load("zerodegress/Index-Translate-2B-MLX-8bit")
prompt = "请将以下中文文本翻译为英语,直接输出翻译结果,不要进行任何解释。\n\n你好,世界!"
print(generate(model, tokenizer, prompt=prompt, max_tokens=1024))

Prompt template: 请将以下{源语言}文本翻译为{目标语言},直接输出翻译结果,不要进行任何解释。 (the source-language slot is left empty for auto-detection). Greedy decoding, thinking disabled.

Files

file bytes sha256
model.safetensors 2,000,043,103 f6d6f9f7dcb8628d2fdeea24b92e50a05075e92eb529d71dd6aea523ec9e62ed

Plus config.json, tokenizer.json, tokenizer_config.json, vocab.json, merges.txt, generation_config.json (copied from upstream), and MANIFEST.json (conversion record).

Attribution

Weights: IndexTeam/Index-Translate-2B (apache-2.0) · family repo. This release only re-encodes the upstream weights; no new rights are claimed.

Downloads last month
36
Safetensors
Model size
2B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zerodegress/Index-Translate-2B-MLX-8bit

Quantized
(15)
this model