# Solomon MLX modifications This port adapts Doccy Pty Ltd’s Apache-2.0 Solomon source at revision `5c0a4a82ddaeca6da2e3013f7045a8196c86957d`. Original modification notices are preserved in `docs/UPSTREAM-MODIFICATIONS.md`. Changes: - Replaced CUDA execution with MLX-VLM Qwen3.5 execution for the declared Qwen3.8 architecture. - Added BF16 shard conversion with FP32 normalization parameters, adapter, trained heads and recurrent states. - Replaced global adapter state with instance-owned state and isolated question cache containers. - Removed vocabulary projection from decision inference. - Added a document-state API, replay recipes, artifact checksums and MLX-specific runtime identities. - Extracted the original prompts, question contract, semantics and evidence utilities into `_vendor`. - Implemented the documented ordering-score product locally because the release omits `scope9.reliability`. - Added download, CUDA parity comparison, benchmark and test tools. New temperature fitting is outside the current scope. No claim is made that this port inherits CUDA calibration or qualification. See `docs/VALIDATION-20260921.md` for measured results and outstanding validation. Adapter-only distribution update (21 September 2026): added a separate verified Hub loader, atomic local assembly and cache reuse; retained the exact historical converter source and FP32 normalization handling. Base weights are downloaded from the pinned Qwen repository. Duplicate adapter/head files under `mlx/bf16` are replaced by the root `adapter/` copies. CUDA parity status is reported separately from conversion.