Instructions to use Accio-Lab/occamy-1.0-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Accio-Lab/occamy-1.0-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Accio-Lab/occamy-1.0-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Accio-Lab/occamy-1.0-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Accio-Lab/occamy-1.0-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Accio-Lab/occamy-1.0-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Accio-Lab/occamy-1.0-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Accio-Lab/occamy-1.0-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Accio-Lab/occamy-1.0-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Accio-Lab/occamy-1.0-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Accio-Lab/occamy-1.0-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Accio-Lab/occamy-1.0-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Occamy-1.0 · MLX 4-bit
Native MLX · group size 64 · 19.51 GB · text only
Candidate release — Mac Metal acceptance is pending. Linux native MLX validation passed. Mac inference, performance and broad model quality remain unverified.
Format
| Property | Value |
|---|---|
| Source | Accio-Lab/occamy-1.0 |
| Source revision | 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8 |
| Quantization | Native affine 4-bit, group size 64 |
| Router / shared-expert gates | 8-bit |
| Weight files | 19,509,024,201 bytes · 19.51 GB · 18.17 GiB |
| Runtime used for Linux checks | mlx 0.32.2, mlx-lm 0.31.3 |
| Inputs | Text only; vision and MTP are not included |
File size is not the unified-memory requirement. Leave room for the operating system, KV cache and runtime buffers. Compare the 3-bit, 4-bit, 6-bit and 8-bit candidates in the MLX collection.
Try with mlx-lm
On Apple Silicon, install the versions used to create this export. The following is a usage recipe awaiting Mac Metal acceptance:
python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3"
from mlx_lm import load, generate
model, tokenizer = load("Accio-Lab/occamy-1.0-MLX-4bit")
messages = [{"role": "user", "content": "Compute 2+2. Answer briefly."}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
response = generate(model, tokenizer, prompt=prompt, max_tokens=128)
print(response)
Start a local API
The stock CLI options below were checked with mlx-lm 0.31.3. Mac Metal acceptance remains pending. HTTP inference checks for this release are recorded only for the new 6-bit and 8-bit exports.
python -m pip install "mlx==0.32.2" "mlx-lm==0.31.3" "transformers==5.8.1"
mlx_lm.server --model Accio-Lab/occamy-1.0-MLX-4bit \
--host 127.0.0.1 --port 8000 \
--chat-template-args '{"enable_thinking":false}'
In a second terminal:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Accio-Lab/occamy-1.0-MLX-4bit","messages":[{"role":"user","content":"Compute 2+2. Answer briefly."}],"temperature":0,"max_tokens":128}'
Client base URL: http://127.0.0.1:8000/v1. Inference runs on your machine.
Conversion and validation
A lossless adapter stacks separate expert weights in numeric expert order before invoking the Qwen3.5 sanitizer exactly once. Quantization and serialization use native APIs; reload uses the stock loader without an adapter.
Passed Linux checks: strict stock reload; complete stored floating-value checks; native dequantization of every quantized row; tokenizer/template comparison; and one bounded cached greedy CPU generation with finite logits. The prompt “Compute 2+2. Answer briefly.” returned 4. This is a limited smoke test, not a quality benchmark or a Mac runtime result.
Validation scope · Artifact hashes · Original model card
License
Apache 2.0, inherited from Occamy-1.0.
Occamy checkpoints
BF16 · GGUF · FP8 · NVFP4 · MLX 8-bit · MLX 6-bit · MLX 4-bit · MLX 3-bit · MTP head
Compare file sizes, validation scope and deployment commands in the checkpoint explorer.
MLX precision family
8bit · 6bit · 5bit · 4bit · 3bit · mxfp8 · mxfp4 · nvfp4
The new 5-bit/MXFP4/MXFP8/NVFP4 cards include their own paired BF16 subset quality checks. Mac Metal acceptance remains pending.
- Downloads last month
- 607
4-bit