botp
/

Solomon / mlx /MODIFICATIONS.md
orz99's picture ArcherHume's picture
Duplicate from DoccyHealth/Solomon
1d2de8a
|
Raw History Blame Contribute Delete
1.62 kB
# Solomon MLX modifications
This port adapts Doccy Pty Ltd’s Apache-2.0 Solomon source at revision
`5c0a4a82ddaeca6da2e3013f7045a8196c86957d`. Original modification notices are preserved in
`docs/UPSTREAM-MODIFICATIONS.md`.
Changes:
- Replaced CUDA execution with MLX-VLM Qwen3.5 execution for the declared Qwen3.8 architecture.
- Added BF16 shard conversion with FP32 normalization parameters, adapter, trained heads and recurrent states.
- Replaced global adapter state with instance-owned state and isolated question cache containers.
- Removed vocabulary projection from decision inference.
- Added a document-state API, replay recipes, artifact checksums and MLX-specific runtime identities.
- Extracted the original prompts, question contract, semantics and evidence utilities into `_vendor`.
- Implemented the documented ordering-score product locally because the release omits `scope9.reliability`.
- Added download, CUDA parity comparison, benchmark and test tools. New temperature fitting is outside the current scope.
No claim is made that this port inherits CUDA calibration or qualification. See
`docs/VALIDATION-20260921.md` for measured results and outstanding validation.
Adapter-only distribution update (21 September 2026): added a separate verified Hub
loader, atomic local assembly and cache reuse; retained the exact historical converter
source and FP32 normalization handling. Base weights are downloaded from the pinned
Qwen repository. Duplicate adapter/head files under `mlx/bf16` are replaced by the
root `adapter/` copies. CUDA parity status is reported separately from conversion.