Lightweight text-only query encoder for webAI-ColVec1.1-4b (CPU-friendly)

#3
by Ryenhails - opened

Hi, and thanks for releasing webAI-ColVec1.1-4b, it's a great model to build on!

We distilled a small text-only query encoder into its late-interaction space (640-d), as part of our ColNanoVDR work (arXiv:2609.34899). It reproduces the model's query token embeddings, so an existing webAI-ColVec1.1-4b page index can be searched with MaxSim unchanged, with no re-indexing and no GPU needed on the query side.

  • nanovdr/ColNanoVDR-Q-Ettin150M-ColVec4B-640-ML: 150M params, 97.4% retention (ViDoRe v1-v3 avg NDCG@5: 71.13 vs. teacher's 73.00)

Training is document-free: only query text and cached teacher query embeddings are used, with no page images or relevance labels. Since it is distilled from your model, the encoder is released under the same webAI Non-Commercial License v1.0, with the full licence text included in the repository.

Would you be open to adding a short note to your model card for users who want to serve queries on CPU? For example:

"Lightweight query encoder: ColNanoVDR-Q-Ettin150M-ColVec4B is a text-only encoder distilled into this model's embedding space, for CPU query encoding against an existing index (non-commercial, same licence as this model)."

Happy to open a PR with this change if that's easier. Thanks again for the great work!

Sign up or log in to comment