Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
RyenhailsΒ 
posted an update 1 day ago
Post
1792
πŸš€ NanoVDR goes multi-vector: meet ColNanoVDR!

Multi-vector VLM retrievers lead visual document retrieval, but every search runs a multi-billion-parameter query encoder. We distill that encoder into a 149M text-only student that queries the teacher's existing page index directly. No re-indexing, and no pages during training.

🧠 How: OTW (Optimal Transport with Learned Weights) aligns the student's query tokens with the teacher's, even though the two tokenize differently (e.g. 17 vs 29 tokens). We prove the alignment cost bounds the MaxSim score gap on every page, so training only needs cached teacher query tokens.

πŸ“Š Five teachers β†’ five 149M students, ViDoRe v3 NDCG@5:
- ColVec1.1-8b: 62.6 β†’ 60.1 (96.0%)
- ColVec1.1-4b: 61.6 β†’ 59.1 (95.8%)
- Vultron-4.5B: 61.0 β†’ 58.3 (95.5%)
- ColQwen3.5-4.5B: 58.7 β†’ 55.1 (93.8%)
- Tomoro-ColQwen3-8B: 59.0 β†’ 54.9 (93.0%)

⚑ 26Γ— faster query encoding on a single CPU thread (87 ms vs 2.3 s)
πŸ’Ύ Matches score distillation while reading 12.6Γ— less cached teacher data

πŸ“„ Paper: ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport (2609.34899)
πŸ€— Checkpoints (all five students):
nanovdr

πŸ’» Code: https://github.com/Ryenhails/NanoVDR
🧩 Single-vector predecessor, NanoVDR: https://arxiv.org/abs/2603.12824

If you already serve one of these teachers, swap in the matching student and keep your index as is. Feedback and upvotes welcome! πŸ™Œ
In this post