Accio

Occamy-1.0 · MLX 4-bit

Native MLX · group size 64 · 19.51 GB · text only

Model collection · Checkpoint explorer · Project · Paper

Candidate release — Mac Metal acceptance is pending. Linux native MLX validation passed. Mac inference, performance and broad model quality remain unverified.

Format

Property Value
Source Accio-Lab/occamy-1.0
Source revision 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8
Quantization Native affine 4-bit, group size 64
Router / shared-expert gates 8-bit
Weight files 19,509,024,201 bytes · 19.51 GB · 18.17 GiB
Runtime used for Linux checks mlx 0.32.2, mlx-lm 0.31.3
Inputs Text only; vision and MTP are not included

File size is not the unified-memory requirement. Leave room for the operating system, KV cache and runtime buffers. Compare the 3-bit, 4-bit, 6-bit and 8-bit candidates in the MLX collection.

Try with mlx-lm

On Apple Silicon, install the versions used to create this export. The following is a usage recipe awaiting Mac Metal acceptance:

python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3"
from mlx_lm import load, generate

model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-4bit")
messages = [{"role": "user", "content": "Compute 2+2. Answer briefly."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
response = generate(model, tokenizer, prompt=prompt, max_tokens=128)
print(response)

Start a local API

The stock CLI options below were checked with mlx-lm 0.31.3. Mac Metal acceptance remains pending. HTTP inference checks for this release are recorded only for the new 6-bit and 8-bit exports.

python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3" "transformers==5.8.1"

mlx_lm.server --model Accio-Lab/occamy-1.0-MLX-4bit \
  --host 127.0.0.1 --port 8000 \
  --chat-template-args '{"enable_thinking":false}'

In a second terminal:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"Accio-Lab/occamy-1.0-MLX-4bit","messages":[{"role":"user","content":"Compute 2+2. Answer briefly."}],"temperature":0,"max_tokens":128}'

Client base URL: http://127.0.0.1:8000/v1. Inference runs on your machine.

Conversion and validation

A lossless adapter stacks separate expert weights in numeric expert order before invoking the Qwen3.5 sanitizer exactly once. Quantization and serialization use native APIs; reload uses the stock loader without an adapter.

Passed Linux checks: strict stock reload; complete stored floating-value checks; native dequantization of every quantized row; tokenizer/template comparison; and one bounded cached greedy CPU generation with finite logits. The prompt “Compute 2+2. Answer briefly.” returned 4. This is a limited smoke test, not a quality benchmark or a Mac runtime result.

Validation scope · Artifact hashes · Original model card

License

Apache 2.0, inherited from Occamy-1.0.

Occamy checkpoints

BF16 · GGUF · FP8 · NVFP4 · MLX 8-bit · MLX 6-bit · MLX 4-bit · MLX 3-bit · MTP head

Compare file sizes, validation scope and deployment commands in the checkpoint explorer.

MLX precision family

8bit · 6bit · 5bit · 4bit · 3bit · mxfp8 · mxfp4 · nvfp4

The new 5-bit/MXFP4/MXFP8/NVFP4 cards include their own paired BF16 subset quality checks. Mac Metal acceptance remains pending.

Downloads last month
607
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Accio-Lab/occamy-1.0-MLX-4bit

Quantized
(31)
this model

Space using Accio-Lab/occamy-1.0-MLX-4bit 1

Collections including Accio-Lab/occamy-1.0-MLX-4bit

Paper for Accio-Lab/occamy-1.0-MLX-4bit