Instructions to use Accio-Lab/occamy-1.0-MLX-5bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Accio-Lab/occamy-1.0-MLX-5bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-5bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Accio-Lab/occamy-1.0-MLX-5bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-5bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Accio-Lab/occamy-1.0-MLX-5bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Accio-Lab/occamy-1.0-MLX-5bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Accio-Lab/occamy-1.0-MLX-5bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-5bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Accio-Lab/occamy-1.0-MLX-5bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Accio-Lab/occamy-1.0-MLX-5bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-5bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Accio-Lab/occamy-1.0-MLX-5bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Accio-Lab/occamy-1.0-MLX-5bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-5bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Accio-Lab/occamy-1.0-MLX-5bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Occamy-1.0 · MLX 5-bit
Native MLX · group size 64 · 23.84 GB · text only
Model collection · MLX collection · Checkpoint explorer · Project · Paper
Candidate release — Mac Metal acceptance is pending. Native Linux MLX artifact, inference, local HTTP and paired held-out quality checks are included. Apple Silicon inference and performance remain unverified.
Format
| Property | Value |
|---|---|
| Source | Accio-Lab/occamy-1.0 |
| Source revision | 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8 |
| Weight quantization | Native MLX affine, 5-bit, group size 64; directly from BF16 |
| Router / shared-expert gates | Affine 8-bit, group size 64 |
| Weight files | 23,838,823,629 bytes · 23.84 GB · 22.202 GiB |
| Export and validation | mlx 0.32.2, mlx-lm 0.31.3, transformers 5.8.1 |
| Inputs | Text only; vision and MTP are separate |
These are MLX weight formats with BF16 inference activations. In particular, MLX NVFP4 is a separate export from the NVIDIA Occamy NVFP4 checkpoint. Weight size does not establish runtime memory use or speed. The MLX family includes 8bit · 6bit · 5bit · 4bit · 3bit · mxfp8 · mxfp4 · nvfp4.
Run a prompt
On Apple Silicon use the pinned versions below; this recipe awaits Metal acceptance.
python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3" "transformers==5.8.1"
python - <<'PY'
from mlx_lm import load, generate
model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-5bit")
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Compute 2+2. Answer briefly."}],
tokenize=False, add_generation_prompt=True, enable_thinking=False,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=128))
PY
Start a local API
python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3" "transformers==5.8.1"
mlx_lm.server --model Accio-Lab/occamy-1.0-MLX-5bit \
--host 127.0.0.1 --port 8000 \
--chat-template-args '{"enable_thinking":false}'
Send a request in a second terminal:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Accio-Lab/occamy-1.0-MLX-5bit","messages":[{"role":"user","content":"Compute 2+2. Answer briefly."}],"temperature":0,"max_tokens":128}'
Use http://127.0.0.1:8000/v1 as an OpenAI-compatible client base URL and Accio-Lab/occamy-1.0-MLX-5bit as the model. Two authored HTTP checks passed on Linux. Mac, tool-call and agent integration acceptance remain unverified. The free CPU explorer supplies commands; inference runs on your hardware.
Conversion and validation
All source payload SHA256 hashes were rechecked. The included lossless adapter stacks 30,720 separate expert tensors in numeric order into 120 groups and invokes the official sanitizer once. Quantization uses stock native MLX APIs, with no requantized input. Exported weights reload directly with stock mlx-lm, without an adapter.
Complete checks cover every stored floating value, native dequantization of every row in all 512 quantized modules, exact tokenizer/template files, strict stock reload and finite full-vocabulary inference logits. Native CPU and CUDA SwitchLinear/MoE kernel probes passed for all four modes. Eight authored cached greedy fixtures covering English/Chinese instructions, arithmetic, JSON and conversation memory passed 8/8. Two stock-server HTTP checks passed 2/2. Per-case outputs and full-file hashes are included.
| Paired native MLX check | Held-out WikiText subset PPL |
|---|---|
| BF16 | 8.3074 |
| MLX 5-bit | 8.3075 |
The same tokenizer, 8,192 token IDs, 16 independent chunks at context 512 and 4,096 scored tokens were used for both. Each chunk has a 256-token unscored prefix and fresh model state. This small test does not establish full benchmark quality. Compare these numbers only within the native MLX scorer; the GGUF results use a separately reported runtime protocol. Reproduction, exact losses and inputs are in quality.
Linux checks used MLX CUDA 12 on NVIDIA B200; conversion used native CPU kernels. Apple Metal, long-context, code, tools, vision and throughput were not tested in this batch. No speed ranking is claimed.
Summary · Artifact and inference results · HTTP results · Conversion receipt · Weight hashes
License and citation
Weights retain Apache 2.0. This checkpoint accompanies Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work. Cite the original report when using the model:
@misc{chen2026occamy10openparetofrontier35b,
title={Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work},
author={Wenhui Chen and Shiwen Cheng and Hao Dong and Chenda Duan and Ruixiang Feng and Zhong Guan and Boqiang Guo and Xueyuan Han and Haojie Hao and Liangmeng Huang and Zhelong Huang and Xinke Kong and Hongyu Li and Jiazheng Li and Junbo Li and Qingchuan Li and Yukun Lian and Chang Liu and Tianyu Liu and Zicheng Liu and Shuyi Ouyang and Yijun Pan and Kunyu Shi and Xiaojun Tang and Bingquan Wang and Kesu Wang and Yuchen Wang and Sibo Wei and Sicong Xie and Xiaoying Xing and Yi Xu and Zhijun Xu and Hongwei Xue and Qingcheng Zeng and Di Zhang and Guannan Zhang and Haochen Zhang and Tianlong Zhang and Tianyu Zhao and Tianyu Zhao and Yanjun Zheng and Jialong Zhu and Zijian Zou},
year={2026},
eprint={2609.11977},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.11977},
}
BF16 · GGUF · APEX GGUF · MLX collection · Explorer
- Downloads last month
- 175
5-bit