botp
/

Solomon / mlx /README.md
orz99's picture ArcherHume's picture
Duplicate from DoccyHealth/Solomon
1d2de8a
|
Raw History Blame Contribute Delete
7.42 kB

Solomon MLX

Full BF16 Solomon v1.1 inference on Apple Silicon. The repository distributes adapters and source code; the Qwen backbone is downloaded separately from its pinned upstream revision and converted locally. No quantization or LoRA merging is performed. Full-model measurements use an M5 Max with 128 GB memory, with about 59.9 GB peak Metal allocation on the focused fixtures. Longer documents and more images require additional memory; 24–32 GB Macs cannot run this profile.

Status: experimental. Focused text and image decisions match CUDA. Complete held-out parity remains pending. Qualification is CUDA parity only: no new calibration or temperature fitting. See validation results.

Install from this repository

Use Python 3.12 or 3.13 on Apple Silicon. Pin the full repository commit shown on Hugging Face, including after any repository history rewrite. The historical Solomon source revision remains provenance; it is not required to be downloadable.

# From a source checkout, enter its mlx/ directory first.
uv sync --frozen --extra dev

# Obtain the current commit once, then retain it for repeatable downloads.
SOLOMON_COMMIT=$(uv run python -c 'from huggingface_hub import HfApi; print(HfApi().model_info("DoccyHealth/Solomon").sha)')
uv run solomon-mlx-hub prepare --revision "$SOLOMON_COMMIT" --output models/quality

Authenticate with hf auth login first if the repository requires access. This package is supplied here as source; it is not claimed to be published on PyPI. For a checkout without downloading model files through Git LFS:

GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/DoccyHealth/Solomon
cd Solomon/mlx

The setup tool explicitly fetches adapter/** and the retained MLX licensing metadata. It downloads Qwen/Qwen3.8-27B at 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, verifies every input against the bundled manifest, and converts one shard at a time. Budget about 112 GB of disk for original and converted weights, plus temporary space and caches. Existing original base files can be reused with --base /path/to/original-qwen.

The output contains backbone/, the unmerged adapter, trained heads, notices and a fresh binding.json. Existing valid outputs are verified and reused; corrupt or incompatible outputs fail without being overwritten. Conversion is atomic and concurrent preparations into the same output are rejected. Interrupted conversions may leave a hidden temporary directory; the final output is never marked ready before verification completes.

For fully offline preparation, provide both downloaded inputs:

uv run solomon-mlx-hub prepare --revision "$SOLOMON_COMMIT" \
  --snapshot /path/to/solomon-snapshot --base /path/to/original-qwen \
  --output models/quality
uv run solomon-mlx-hub verify models/quality

CPU conversion is the default. --device gpu selects Metal conversion on a Mac. Linux CPU conversion is also supported by the existing converter and can use uv sync --frozen --extra cloud with MLX's CPU backend. Linux conversion does not run Apple Metal inference or establish CUDA parity.

Converter and provenance

The exact converter is src/solomon_mlx/prepare.py. The runtime Python sources are unchanged in behaviour. The only edits made for this release rename the runtime identity's source_contract field to solomon-v1 and adjust comments, docstrings and one error message, so their combined converter_code_sha256 is bbcae17fc1c35db80a79d5865133a42ef9a1b0cf342fff71949972e69cce43ec. It is computed by solomon_mlx.artifacts.code_identity() from the sorted mapping of relative Python paths to SHA-256 values. The new download/assembly wrapper is in the separate solomon_mlx_hub package.

The converter retains large weights in BF16 and promotes normalization weights, A_log and dt_bias to FP32 before applying upstream normalization offsets. It validates tensor names and shapes with MLX-VLM's Qwen3.5 implementation. The original FP32 adapter and ten trained heads are copied without changes. The runtime applies the adapter only to question tokens, with its explicit 2.0 scale.

bf16/conversion.json and bf16/binding.json describe the historical cloud conversion; paths inside them describe the original local model layout. They are provenance, not a manifest of files currently present on the Hub, and they record the converter hash of that historical conversion, 648e440cface0838f2dcc8d89b3ab172d97f4ffc1f88fe0b7cbbe3ca73b8a575, which predates the documentation edits described above. The root adapter files have the same hashes as their removed duplicates. Newly converted safetensors may serialize differently across CPU and Metal, so new outputs get actual output checksums and their own runtime binding. Historical CUDA or MLX qualification identities are never reused for a different artifact binding.

Python API

from solomon_mlx import Solomon

model = Solomon.load("models/quality", profile="quality")
with model.prefill("Rookwood Ltd holds a current certification.") as state:
    result = model.decide(
        state=state,
        questions={"certified": {
            "type": "noul",
            "instructions": "Does Rookwood Ltd hold a current certification?",
        }},
        evidence="support",
    )
    print(result["answers"]["certified"])

Or use solomon_mlx_hub.load("models/quality", revision=COMMIT) to prepare and load in one call. revision must be a full 40-character commit SHA. Cached valid outputs are reused without network access and retain their original binding.

Documents accept text, structured JSON objects, or ordered {"text": ...} and {"image": local_path} parts. Image features are computed once per document. There is no PDF renderer or OCR. Question forms preserve Boolean, entity and multilabel noul, single choice, and ordered score semantics; candidate order is retained. Missing/conflicting facts collapse before temperature application. The API defaults to T=1. No new temperatures are fitted by setup or parity checks.

Evidence levels are none, support, sufficiency, and removal. Text spans use exact code-point offsets. Sufficiency and removal re-encode the relevant source; they do not establish causal faithfulness. Image evidence requires page_selector=; otherwise the result reports unsupported_page_selector. States belong to one model instance and support close(), save(path), and model.replay(path). Replay persists a checksummed source recipe and recomputes caches.

Tests

uv run pytest -q
uv run ruff check src tests scripts

Tests use synthetic small models to cover conversion precision, CPU/Metal tensor agreement, cache isolation, text/image processing, answer semantics, evidence, artifact corruption, and adapter-only setup. These tests do not replace trained model parity measurements. The included parity checker compares saved CUDA and MLX scores using the same frozen temperatures exactly once, reports probability drift separately, and cannot pass a complete-panel gate from partial results. Evaluation inputs and private infrastructure configuration are not distributed.

Preserve LICENSE, NOTICE, MODIFICATIONS.md, and upstream model notices with permitted copies. Model downloads and conversion do not alter repository visibility.