Jeeves-9B MLX — 4bit

Unofficial CPU-converted MLX packaging of PostHog/jeeves, published by cowWhySo. This is not a PostHog or Apple release, a newly trained model, or a benchmark-certified port.

Original project: PostHog/jeeves on GitHub
Original model: PostHog/jeeves on Hugging Face

Validation: converted_and_small_cpu_diagnostics_passed_not_benchmarked. Read validation.json before relying on the package. Apple Metal execution, task accuracy, post-quantization calibration, and representative latency remain unmeasured. Upstream benchmark scores and CUDA speed claims do not transfer to this conversion.

Available packages

One repository; one complete package per branch. main contains 4-bit, not BF16. Other variants are branches, not subfolders. Downloading one revision does not download the others.

Variant Revision Original package size on Drive
4bit main 4.73 GiB
8bit 8bit 8.90 GiB
bf16 bf16 16.71 GiB

Sizes are actual original package bytes divided by 2^30 (GiB), including auxiliary files. Publishing adds small documentation/manifest files. These are disk sizes, not measured peak RAM or unified-memory requirements. This README describes the main revision.

What is preserved

The Qwen3.5-9B backbone uses 4-bit MLX affine quantization, group size 64. The original pointer head is retained as head.pt and exported without changing its FP32 values as head.safetensors. The released fitted temperature is 1.85892808437347, retained in export.json; it has not been refitted for quantization. Exact-format helpers, the tokenizer, original licenses, source records, and diagnostic fixtures accompany the weights.

Jeeves uses a custom decision head and prompt/readout format, not an ordinary text-generation pipeline. Use the included run_jeeves.py for decisions; mlx_lm.generate alone does not implement Jeeves scoring. See BUILD.md for build details and provenance/original_package_README.md for the original conversion notes.

Download and try on Apple Silicon

These are instructions for an unmeasured Metal path, not a claim that this release has been tested on a Mac. Review the included Python source before running it. Each variant should have its own local directory; do not overlay variants into one directory.

python3 -m venv .venv
source .venv/bin/activate
python -m pip install "huggingface_hub==1.33.0"
hf download cowWhySo/jeeves-mlx --revision main --local-dir jeeves-mlx-4bit
cd jeeves-mlx-4bit
python -m pip install -r requirements-mac.txt
python run_jeeves.py --model . --request example_request.json --device gpu --think-tokens 0

To select a different variant, replace the revision and directory with main / jeeves-mlx-4bit, 8bit / jeeves-mlx-8bit, or bf16 / jeeves-mlx-bf16. Pin a release commit SHA instead of a branch for reproducible downstream experiments.

The runner accepts state and a questions mapping with noul, choice, or score types. It returns option probabilities under its own results schema; it is not wire-compatible with the upstream /v1/systemone HTTP service.

An optional, unvalidated greedy-thinking path exists:

python run_jeeves.py --model . --request example_request.json --device gpu --think-tokens 128

The reference runner processes questions independently, uses a separate cache per question, and re-prefills the full sequence for decision readout after optional reasoning. It does not port the speculative diffusion drafters, CUDA graphs, FP8 serving, parallel-question engine, or adaptive no-thinking threshold.

Evidence and limits

The stored format/head checks passed for 4 format cases. The synthetic-hidden-state pointer-head comparison reported maximum probability error 0; it is not a backbone comparison. The stored no-thinking CPU smoke test contains 1 case(s). The BF16 backbone comparison is explicitly a small-fixture diagnostic with heuristic tolerances, not general equivalence.

On the recorded 1-case diagnostic, the maximum option-probability difference from MLX BF16 was 0.016297996. This is not an accuracy or calibration estimate.

No task accuracy, calibration benchmark, Metal speed, quantized quality ranking, or validated reasoning performance is claimed. CPU smoke elapsed times in the evidence are diagnostics, not representative serving benchmarks. Retaining the original temperature is not evidence of calibration on a different numerical path.

Build provenance

PACKAGE_MANIFEST.json hashes the published payload except itself. provenance/source_PACKAGE_MANIFEST.json is the unmodified original manifest; its paths describe the original pre-publication package, not the newly documented payload. RELEASE.json distinguishes conversion evidence from packaging checks. Original validation.json is unchanged.

License and attribution

Model weights follow the original Apache-2.0 license retained in LICENSE. Upstream Jeeves code is MIT licensed; its notice is retained in LICENSE_JEEVES_CODE. Respect any additional notices in the included files. Credits for the original model, training, calibration, and serving research belong to PostHog and the upstream authors; this repository provides an unofficial MLX conversion/packaging.

Downloads last month
15
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cowWhySo/jeeves-mlx

Finetuned
Qwen/Qwen3.5-9B
Finetuned
PostHog/jeeves
Quantized
(3)
this model