# Solomon MLX Full BF16 Solomon v1.1 inference on Apple Silicon. The repository distributes adapters and source code; the Qwen backbone is downloaded separately from its pinned upstream revision and converted locally. No quantization or LoRA merging is performed. Full-model measurements use an M5 Max with 128 GB memory, with about 59.9 GB peak Metal allocation on the focused fixtures. Longer documents and more images require additional memory; 24–32 GB Macs cannot run this profile. **Status: experimental.** Focused text and image decisions match CUDA. Complete held-out parity remains pending. Qualification is CUDA parity only: no new calibration or temperature fitting. See [validation results](docs/VALIDATION-20260921.md). ## Install from this repository Use Python 3.12 or 3.13 on Apple Silicon. Pin the full repository commit shown on Hugging Face, including after any repository history rewrite. The historical Solomon source revision remains provenance; it is not required to be downloadable. ```sh # From a source checkout, enter its mlx/ directory first. uv sync --frozen --extra dev # Obtain the current commit once, then retain it for repeatable downloads. SOLOMON_COMMIT=$(uv run python -c 'from huggingface_hub import HfApi; print(HfApi().model_info("DoccyHealth/Solomon").sha)') uv run solomon-mlx-hub prepare --revision "$SOLOMON_COMMIT" --output models/quality ``` Authenticate with `hf auth login` first if the repository requires access. This package is supplied here as source; it is not claimed to be published on PyPI. For a checkout without downloading model files through Git LFS: ```sh GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/DoccyHealth/Solomon cd Solomon/mlx ``` The setup tool explicitly fetches `adapter/**` and the retained MLX licensing metadata. It downloads `Qwen/Qwen3.8-27B` at `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0`, verifies every input against the bundled manifest, and converts one shard at a time. Budget about 112 GB of disk for original and converted weights, plus temporary space and caches. Existing original base files can be reused with `--base /path/to/original-qwen`. The output contains `backbone/`, the unmerged adapter, trained heads, notices and a fresh `binding.json`. Existing valid outputs are verified and reused; corrupt or incompatible outputs fail without being overwritten. Conversion is atomic and concurrent preparations into the same output are rejected. Interrupted conversions may leave a hidden temporary directory; the final output is never marked ready before verification completes. For fully offline preparation, provide both downloaded inputs: ```sh uv run solomon-mlx-hub prepare --revision "$SOLOMON_COMMIT" \ --snapshot /path/to/solomon-snapshot --base /path/to/original-qwen \ --output models/quality uv run solomon-mlx-hub verify models/quality ``` CPU conversion is the default. `--device gpu` selects Metal conversion on a Mac. Linux CPU conversion is also supported by the existing converter and can use `uv sync --frozen --extra cloud` with MLX's CPU backend. Linux conversion does not run Apple Metal inference or establish CUDA parity. ## Converter and provenance The exact converter is [src/solomon_mlx/prepare.py](src/solomon_mlx/prepare.py). The runtime Python sources are unchanged in behaviour. The only edits made for this release rename the runtime identity's `source_contract` field to `solomon-v1` and adjust comments, docstrings and one error message, so their combined `converter_code_sha256` is `bbcae17fc1c35db80a79d5865133a42ef9a1b0cf342fff71949972e69cce43ec`. It is computed by `solomon_mlx.artifacts.code_identity()` from the sorted mapping of relative Python paths to SHA-256 values. The new download/assembly wrapper is in the separate `solomon_mlx_hub` package. The converter retains large weights in BF16 and promotes normalization weights, `A_log` and `dt_bias` to FP32 **before** applying upstream normalization offsets. It validates tensor names and shapes with MLX-VLM's Qwen3.5 implementation. The original FP32 adapter and ten trained heads are copied without changes. The runtime applies the adapter only to question tokens, with its explicit 2.0 scale. `bf16/conversion.json` and `bf16/binding.json` describe the historical cloud conversion; paths inside them describe the original local model layout. They are provenance, not a manifest of files currently present on the Hub, and they record the converter hash of that historical conversion, `648e440cface0838f2dcc8d89b3ab172d97f4ffc1f88fe0b7cbbe3ca73b8a575`, which predates the documentation edits described above. The root adapter files have the same hashes as their removed duplicates. Newly converted safetensors may serialize differently across CPU and Metal, so new outputs get actual output checksums and their own runtime binding. Historical CUDA or MLX qualification identities are never reused for a different artifact binding. ## Python API ```python from solomon_mlx import Solomon model = Solomon.load("models/quality", profile="quality") with model.prefill("Rookwood Ltd holds a current certification.") as state: result = model.decide( state=state, questions={"certified": { "type": "noul", "instructions": "Does Rookwood Ltd hold a current certification?", }}, evidence="support", ) print(result["answers"]["certified"]) ``` Or use `solomon_mlx_hub.load("models/quality", revision=COMMIT)` to prepare and load in one call. `revision` must be a full 40-character commit SHA. Cached valid outputs are reused without network access and retain their original binding. Documents accept text, structured JSON objects, or ordered `{"text": ...}` and `{"image": local_path}` parts. Image features are computed once per document. There is no PDF renderer or OCR. Question forms preserve Boolean, entity and multilabel `noul`, single `choice`, and ordered `score` semantics; candidate order is retained. Missing/conflicting facts collapse before temperature application. The API defaults to T=1. No new temperatures are fitted by setup or parity checks. Evidence levels are `none`, `support`, `sufficiency`, and `removal`. Text spans use exact code-point offsets. Sufficiency and removal re-encode the relevant source; they do not establish causal faithfulness. Image evidence requires `page_selector=`; otherwise the result reports `unsupported_page_selector`. States belong to one model instance and support `close()`, `save(path)`, and `model.replay(path)`. Replay persists a checksummed source recipe and recomputes caches. ## Tests ```sh uv run pytest -q uv run ruff check src tests scripts ``` Tests use synthetic small models to cover conversion precision, CPU/Metal tensor agreement, cache isolation, text/image processing, answer semantics, evidence, artifact corruption, and adapter-only setup. These tests do not replace trained model parity measurements. The included parity checker compares saved CUDA and MLX scores using the same frozen temperatures exactly once, reports probability drift separately, and cannot pass a complete-panel gate from partial results. Evaluation inputs and private infrastructure configuration are not distributed. Preserve `LICENSE`, `NOTICE`, `MODIFICATIONS.md`, and upstream model notices with permitted copies. Model downloads and conversion do not alter repository visibility.