Thomson-1.0-Small — MLX 4-bit Quantization

MLX format conversion of thomsonreuters/Thomson-1.0-Small, a 35B parameter Mixture-of-Experts VLM with hybrid linear attention.

Base architecture: Qwen3.5-35B-A3B (hybrid attention MoE, 256 experts, 8 active per token) Vision tower: Qwen3.5 Vision Encoder, kept at fp16 (not quantized)

Model Details

Property Value
Bits 4
Group size 64
Effective BPW 4.65
Size ~20.4 GB
Shards 4
Quantization Affine, RTN (round-to-nearest)
Vision tower fp16 (preserved)
Dtype bfloat16
Context length 262,144

Quickstart

pip install -U mlx-vlm

python3 -m mlx_vlm.generate \
  --model hermitdave/Thomson-1.0-Small-MLX-4bit \
  --prompt "Summarize the key points of this document." \
  --max-tokens 512 --temp 1.0 --top-p 0.95

For deterministic outputs, use temp=0. For complex reasoning tasks, increase max-tokens.

Conversion Details

  • Tool: mlx-vlm v0.6.17 convert (lazy mmap mode)
  • Quantization: Uniform affine, RTN, no mixed-predicate
  • Vision tower: Preserved at fp16 via skip_multimodal_module predicate
  • Dtype: bfloat16
  • Hardware: Apple M3 Max 64 GB
  • Conversion script: convert_thomson_4b_6b.py

Attribution

Upstream model: thomsonreuters/Thomson-1.0-Small by Thomson Reuters.

Thomson-1.0 stem: tri-fair-lab/Snowdon1.1-Small by Tair Lab.

Conversion: Hermes Agent (Nous Research) using mlx-vlm v0.6.17 on Apple Silicon.

Technical report: "Thomson: Continual Learning of Frontier Models for SovereignAI" (arXiv:2608.27147).

Other Formats

Also available: 6-bit variant

License

Same as upstream: Qwen3.6 Community License. See upstream repo for full terms.

Disclaimer

Quantized models may exhibit slightly different behavior compared to the original BF16 checkpoint. Please validate for your specific use case.

Downloads last month
115
Safetensors
Model size
35B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hermitdave/Thomson-1.0-Small-MLX-4bit

Quantized
(12)
this model

Paper for hermitdave/Thomson-1.0-Small-MLX-4bit