Accio

Occamy-1.0 · MLX 5-bit

Native MLX · group size 64 · 23.84 GB · text only

Model collection · MLX collection · Checkpoint explorer · Project · Paper

Candidate release — Mac Metal acceptance is pending. Native Linux MLX artifact, inference, local HTTP and paired held-out quality checks are included. Apple Silicon inference and performance remain unverified.

Format

Property Value
Source Accio-Lab/occamy-1.0
Source revision 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8
Weight quantization Native MLX affine, 5-bit, group size 64; directly from BF16
Router / shared-expert gates Affine 8-bit, group size 64
Weight files 23,838,823,629 bytes · 23.84 GB · 22.202 GiB
Export and validation mlx 0.32.2, mlx-lm 0.31.3, transformers 5.8.1
Inputs Text only; vision and MTP are separate

These are MLX weight formats with BF16 inference activations. In particular, MLX NVFP4 is a separate export from the NVIDIA Occamy NVFP4 checkpoint. Weight size does not establish runtime memory use or speed. The MLX family includes 8bit · 6bit · 5bit · 4bit · 3bit · mxfp8 · mxfp4 · nvfp4.

Run a prompt

On Apple Silicon use the pinned versions below; this recipe awaits Metal acceptance.

python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3" "transformers==5.8.1"

python - <<'PY'
from mlx_lm import load, generate
model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-5bit")
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Compute 2+2. Answer briefly."}],
    tokenize=False, add_generation_prompt=True, enable_thinking=False,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=128))
PY

Start a local API

python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3" "transformers==5.8.1"

mlx_lm.server --model Accio-Lab/occamy-1.0-MLX-5bit \
  --host 127.0.0.1 --port 8000 \
  --chat-template-args '{"enable_thinking":false}'

Send a request in a second terminal:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"Accio-Lab/occamy-1.0-MLX-5bit","messages":[{"role":"user","content":"Compute 2+2. Answer briefly."}],"temperature":0,"max_tokens":128}'

Use http://127.0.0.1:8000/v1 as an OpenAI-compatible client base URL and Accio-Lab/occamy-1.0-MLX-5bit as the model. Two authored HTTP checks passed on Linux. Mac, tool-call and agent integration acceptance remain unverified. The free CPU explorer supplies commands; inference runs on your hardware.

Conversion and validation

All source payload SHA256 hashes were rechecked. The included lossless adapter stacks 30,720 separate expert tensors in numeric order into 120 groups and invokes the official sanitizer once. Quantization uses stock native MLX APIs, with no requantized input. Exported weights reload directly with stock mlx-lm, without an adapter.

Complete checks cover every stored floating value, native dequantization of every row in all 512 quantized modules, exact tokenizer/template files, strict stock reload and finite full-vocabulary inference logits. Native CPU and CUDA SwitchLinear/MoE kernel probes passed for all four modes. Eight authored cached greedy fixtures covering English/Chinese instructions, arithmetic, JSON and conversation memory passed 8/8. Two stock-server HTTP checks passed 2/2. Per-case outputs and full-file hashes are included.

Paired native MLX check Held-out WikiText subset PPL
BF16 8.3074
MLX 5-bit 8.3075

The same tokenizer, 8,192 token IDs, 16 independent chunks at context 512 and 4,096 scored tokens were used for both. Each chunk has a 256-token unscored prefix and fresh model state. This small test does not establish full benchmark quality. Compare these numbers only within the native MLX scorer; the GGUF results use a separately reported runtime protocol. Reproduction, exact losses and inputs are in quality.

Linux checks used MLX CUDA 12 on NVIDIA B200; conversion used native CPU kernels. Apple Metal, long-context, code, tools, vision and throughput were not tested in this batch. No speed ranking is claimed.

Summary · Artifact and inference results · HTTP results · Conversion receipt · Weight hashes

License and citation

Weights retain Apache 2.0. This checkpoint accompanies Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work. Cite the original report when using the model:

@misc{chen2026occamy10openparetofrontier35b,
      title={Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work}, 
      author={Wenhui Chen and Shiwen Cheng and Hao Dong and Chenda Duan and Ruixiang Feng and Zhong Guan and Boqiang Guo and Xueyuan Han and Haojie Hao and Liangmeng Huang and Zhelong Huang and Xinke Kong and Hongyu Li and Jiazheng Li and Junbo Li and Qingchuan Li and Yukun Lian and Chang Liu and Tianyu Liu and Zicheng Liu and Shuyi Ouyang and Yijun Pan and Kunyu Shi and Xiaojun Tang and Bingquan Wang and Kesu Wang and Yuchen Wang and Sibo Wei and Sicong Xie and Xiaoying Xing and Yi Xu and Zhijun Xu and Hongwei Xue and Qingcheng Zeng and Di Zhang and Guannan Zhang and Haochen Zhang and Tianlong Zhang and Tianyu Zhao and Tianyu Zhao and Yanjun Zheng and Jialong Zhu and Zijian Zou},
      year={2026},
      eprint={2609.11977},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.11977}, 
}

BF16 · GGUF · APEX GGUF · MLX collection · Explorer

Downloads last month
175
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Accio-Lab/occamy-1.0-MLX-5bit

Quantized
(31)
this model

Space using Accio-Lab/occamy-1.0-MLX-5bit 1

Collections including Accio-Lab/occamy-1.0-MLX-5bit

Paper for Accio-Lab/occamy-1.0-MLX-5bit