botp
/

Solomon / mlx /MODIFICATIONS.md
orz99's picture ArcherHume's picture
Duplicate from DoccyHealth/Solomon
1d2de8a
|
Raw History Blame Contribute Delete
1.62 kB

Solomon MLX modifications

This port adapts Doccy Pty Ltd’s Apache-2.0 Solomon source at revision 5c0a4a82ddaeca6da2e3013f7045a8196c86957d. Original modification notices are preserved in docs/UPSTREAM-MODIFICATIONS.md.

Changes:

  • Replaced CUDA execution with MLX-VLM Qwen3.5 execution for the declared Qwen3.8 architecture.
  • Added BF16 shard conversion with FP32 normalization parameters, adapter, trained heads and recurrent states.
  • Replaced global adapter state with instance-owned state and isolated question cache containers.
  • Removed vocabulary projection from decision inference.
  • Added a document-state API, replay recipes, artifact checksums and MLX-specific runtime identities.
  • Extracted the original prompts, question contract, semantics and evidence utilities into _vendor.
  • Implemented the documented ordering-score product locally because the release omits scope9.reliability.
  • Added download, CUDA parity comparison, benchmark and test tools. New temperature fitting is outside the current scope.

No claim is made that this port inherits CUDA calibration or qualification. See docs/VALIDATION-20260921.md for measured results and outstanding validation.

Adapter-only distribution update (21 September 2026): added a separate verified Hub loader, atomic local assembly and cache reuse; retained the exact historical converter source and FP32 normalization handling. Base weights are downloaded from the pinned Qwen repository. Duplicate adapter/head files under mlx/bf16 are replaced by the root adapter/ copies. CUDA parity status is reported separately from conversion.