Instructions to use webAI-Official/webAI-ColVec1.1-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use webAI-Official/webAI-ColVec1.1-4b with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("webAI-Official/webAI-ColVec1.1-4b", trust_remote_code=True) model = AutoModel.from_pretrained("webAI-Official/webAI-ColVec1.1-4b", trust_remote_code=True, device_map="auto") - sentence-transformers
How to use webAI-Official/webAI-ColVec1.1-4b with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("webAI-Official/webAI-ColVec1.1-4b", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Lightweight text-only query encoder for webAI-ColVec1.1-4b (CPU-friendly)
Hi, and thanks for releasing webAI-ColVec1.1-4b, it's a great model to build on!
We distilled a small text-only query encoder into its late-interaction space (640-d), as part of our ColNanoVDR work (arXiv:2609.34899). It reproduces the model's query token embeddings, so an existing webAI-ColVec1.1-4b page index can be searched with MaxSim unchanged, with no re-indexing and no GPU needed on the query side.
- nanovdr/ColNanoVDR-Q-Ettin150M-ColVec4B-640-ML: 150M params, 97.4% retention (ViDoRe v1-v3 avg NDCG@5: 71.13 vs. teacher's 73.00)
Training is document-free: only query text and cached teacher query embeddings are used, with no page images or relevance labels. Since it is distilled from your model, the encoder is released under the same webAI Non-Commercial License v1.0, with the full licence text included in the repository.
Would you be open to adding a short note to your model card for users who want to serve queries on CPU? For example:
"Lightweight query encoder: ColNanoVDR-Q-Ettin150M-ColVec4B is a text-only encoder distilled into this model's embedding space, for CPU query encoding against an existing index (non-commercial, same licence as this model)."
Happy to open a PR with this change if that's easier. Thanks again for the great work!