botp
/

Solomon / mlx /README.md
orz99's picture ArcherHume's picture
Duplicate from DoccyHealth/Solomon
1d2de8a
|
Raw History Blame Contribute Delete
7.42 kB
# Solomon MLX
Full BF16 Solomon v1.1 inference on Apple Silicon. The repository distributes
adapters and source code; the Qwen backbone is downloaded separately from its
pinned upstream revision and converted locally. No quantization or LoRA merging
is performed. Full-model measurements use an M5 Max with 128 GB memory, with
about 59.9 GB peak Metal allocation on the focused fixtures. Longer documents
and more images require additional memory; 24–32 GB Macs cannot run this profile.
**Status: experimental.** Focused text and image decisions match CUDA. Complete
held-out parity remains pending. Qualification is CUDA parity only: no new
calibration or temperature fitting. See [validation results](docs/VALIDATION-20260921.md).
## Install from this repository
Use Python 3.12 or 3.13 on Apple Silicon. Pin the full repository commit shown on
Hugging Face, including after any repository history rewrite. The historical
Solomon source revision remains provenance; it is not required to be downloadable.
```sh
# From a source checkout, enter its mlx/ directory first.
uv sync --frozen --extra dev
# Obtain the current commit once, then retain it for repeatable downloads.
SOLOMON_COMMIT=$(uv run python -c 'from huggingface_hub import HfApi; print(HfApi().model_info("DoccyHealth/Solomon").sha)')
uv run solomon-mlx-hub prepare --revision "$SOLOMON_COMMIT" --output models/quality
```
Authenticate with `hf auth login` first if the repository requires access. This
package is supplied here as source; it is not claimed to be published on PyPI.
For a checkout without downloading model files through Git LFS:
```sh
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/DoccyHealth/Solomon
cd Solomon/mlx
```
The setup tool explicitly fetches `adapter/**` and the retained MLX licensing
metadata. It downloads `Qwen/Qwen3.8-27B` at
`1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0`, verifies every input against the
bundled manifest, and converts one shard at a time. Budget about 112 GB of disk
for original and converted weights, plus temporary space and caches. Existing
original base files can be reused with `--base /path/to/original-qwen`.
The output contains `backbone/`, the unmerged adapter, trained heads, notices
and a fresh `binding.json`. Existing valid outputs are verified and reused;
corrupt or incompatible outputs fail without being overwritten. Conversion is
atomic and concurrent preparations into the same output are rejected. Interrupted
conversions may leave a hidden temporary directory; the final output is never
marked ready before verification completes.
For fully offline preparation, provide both downloaded inputs:
```sh
uv run solomon-mlx-hub prepare --revision "$SOLOMON_COMMIT" \
--snapshot /path/to/solomon-snapshot --base /path/to/original-qwen \
--output models/quality
uv run solomon-mlx-hub verify models/quality
```
CPU conversion is the default. `--device gpu` selects Metal conversion on a Mac.
Linux CPU conversion is also supported by the existing converter and can use
`uv sync --frozen --extra cloud` with MLX's CPU backend. Linux conversion does
not run Apple Metal inference or establish CUDA parity.
## Converter and provenance
The exact converter is [src/solomon_mlx/prepare.py](src/solomon_mlx/prepare.py).
The runtime Python sources are unchanged in behaviour. The only edits made for
this release rename the runtime identity's `source_contract` field to `solomon-v1`
and adjust comments, docstrings and one error message, so their combined
`converter_code_sha256` is
`bbcae17fc1c35db80a79d5865133a42ef9a1b0cf342fff71949972e69cce43ec`.
It is computed by `solomon_mlx.artifacts.code_identity()` from the sorted mapping
of relative Python paths to SHA-256 values. The new download/assembly wrapper is
in the separate `solomon_mlx_hub` package.
The converter retains large weights in BF16 and promotes normalization weights,
`A_log` and `dt_bias` to FP32 **before** applying upstream normalization offsets.
It validates tensor names and shapes with MLX-VLM's Qwen3.5 implementation. The
original FP32 adapter and ten trained heads are copied without changes. The
runtime applies the adapter only to question tokens, with its explicit 2.0 scale.
`bf16/conversion.json` and `bf16/binding.json` describe the historical cloud
conversion; paths inside them describe the original local model layout. They
are provenance, not a manifest of files currently present on the Hub, and they
record the converter hash of that historical conversion, `648e440cface0838f2dcc8d89b3ab172d97f4ffc1f88fe0b7cbbe3ca73b8a575`,
which predates the documentation edits described above. The root
adapter files have the same hashes as their removed duplicates. Newly converted
safetensors may serialize differently across CPU and Metal, so new outputs get
actual output checksums and their own runtime binding. Historical CUDA or MLX
qualification identities are never reused for a different artifact binding.
## Python API
```python
from solomon_mlx import Solomon
model = Solomon.load("models/quality", profile="quality")
with model.prefill("Rookwood Ltd holds a current certification.") as state:
result = model.decide(
state=state,
questions={"certified": {
"type": "noul",
"instructions": "Does Rookwood Ltd hold a current certification?",
}},
evidence="support",
)
print(result["answers"]["certified"])
```
Or use `solomon_mlx_hub.load("models/quality", revision=COMMIT)` to prepare and
load in one call. `revision` must be a full 40-character commit SHA. Cached valid
outputs are reused without network access and retain their original binding.
Documents accept text, structured JSON objects, or ordered `{"text": ...}` and
`{"image": local_path}` parts. Image features are computed once per document.
There is no PDF renderer or OCR. Question forms preserve Boolean, entity and
multilabel `noul`, single `choice`, and ordered `score` semantics; candidate order
is retained. Missing/conflicting facts collapse before temperature application.
The API defaults to T=1. No new temperatures are fitted by setup or parity checks.
Evidence levels are `none`, `support`, `sufficiency`, and `removal`. Text spans use
exact code-point offsets. Sufficiency and removal re-encode the relevant source;
they do not establish causal faithfulness. Image evidence requires `page_selector=`;
otherwise the result reports `unsupported_page_selector`. States belong to one
model instance and support `close()`, `save(path)`, and `model.replay(path)`.
Replay persists a checksummed source recipe and recomputes caches.
## Tests
```sh
uv run pytest -q
uv run ruff check src tests scripts
```
Tests use synthetic small models to cover conversion precision, CPU/Metal tensor
agreement, cache isolation, text/image processing, answer semantics, evidence,
artifact corruption, and adapter-only setup. These tests do not replace trained
model parity measurements. The included parity checker compares saved CUDA and
MLX scores using the same frozen temperatures exactly once, reports probability
drift separately, and cannot pass a complete-panel gate from partial results.
Evaluation inputs and private infrastructure configuration are not distributed.
Preserve `LICENSE`, `NOTICE`, `MODIFICATIONS.md`, and upstream model notices with
permitted copies. Model downloads and conversion do not alter repository visibility.