colSmol-256M-6bit

An MLX copy of vidore/colSmol-256M, a page retriever: it turns a page picture, or a text query, into one 128-number vector per token, and pages are ranked against a query by late interaction (MaxSim).

It does not chat. Use it to find the right page in a set of documents.

How this copy was made

  1. The LoRA adapter in vidore/colSmol-256M was merged into the weights of its base, vidore/ColSmolVLM-Instruct-256M-base.
  2. The config's model_type was set to colidefics3, the type mlx-vlm loads.
  3. The merged weights were quantized to 6 bits (affine, group size 64) with mlx-vlm 0.7.6.

Nothing was retrained. The merge was checked against the original PyTorch weights: MaxSim scores agree within 0.6%. Six-bit rounding moves the vectors by about 0.006 on average.

Use

With mlx-vlm in Python, load it as any colidefics3 checkpoint.

In Swift, Different-Productions/mlx-swift-lm loads it through MLXEmbedders (ColIdefics3Model, ColIdefics3Processor), and matches mlx-vlm on the same files: the same token ids and ranking, scores within 0.7% on PNG pages.

License and credit

MIT, as the original. The model is the work of the ViDoRe team at Illuin Technology; see the original card for training data, benchmarks and how to cite it. Its backbone, SmolVLM-256M-Instruct by Hugging Face, is Apache 2.0.

This repository changes the original in the three ways listed above and no others.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DifferentProductions/colSmol-256M-6bit