Post
1792
π NanoVDR goes multi-vector: meet ColNanoVDR!
Multi-vector VLM retrievers lead visual document retrieval, but every search runs a multi-billion-parameter query encoder. We distill that encoder into a 149M text-only student that queries the teacher's existing page index directly. No re-indexing, and no pages during training.
π§ How: OTW (Optimal Transport with Learned Weights) aligns the student's query tokens with the teacher's, even though the two tokenize differently (e.g. 17 vs 29 tokens). We prove the alignment cost bounds the MaxSim score gap on every page, so training only needs cached teacher query tokens.
π Five teachers β five 149M students, ViDoRe v3 NDCG@5:
- ColVec1.1-8b: 62.6 β 60.1 (96.0%)
- ColVec1.1-4b: 61.6 β 59.1 (95.8%)
- Vultron-4.5B: 61.0 β 58.3 (95.5%)
- ColQwen3.5-4.5B: 58.7 β 55.1 (93.8%)
- Tomoro-ColQwen3-8B: 59.0 β 54.9 (93.0%)
β‘ 26Γ faster query encoding on a single CPU thread (87 ms vs 2.3 s)
πΎ Matches score distillation while reading 12.6Γ less cached teacher data
π Paper: ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport (2609.34899)
π€ Checkpoints (all five students):
nanovdr
π» Code: https://github.com/Ryenhails/NanoVDR
π§© Single-vector predecessor, NanoVDR: https://arxiv.org/abs/2603.12824
If you already serve one of these teachers, swap in the matching student and keep your index as is. Feedback and upvotes welcome! π
Multi-vector VLM retrievers lead visual document retrieval, but every search runs a multi-billion-parameter query encoder. We distill that encoder into a 149M text-only student that queries the teacher's existing page index directly. No re-indexing, and no pages during training.
π§ How: OTW (Optimal Transport with Learned Weights) aligns the student's query tokens with the teacher's, even though the two tokenize differently (e.g. 17 vs 29 tokens). We prove the alignment cost bounds the MaxSim score gap on every page, so training only needs cached teacher query tokens.
π Five teachers β five 149M students, ViDoRe v3 NDCG@5:
- ColVec1.1-8b: 62.6 β 60.1 (96.0%)
- ColVec1.1-4b: 61.6 β 59.1 (95.8%)
- Vultron-4.5B: 61.0 β 58.3 (95.5%)
- ColQwen3.5-4.5B: 58.7 β 55.1 (93.8%)
- Tomoro-ColQwen3-8B: 59.0 β 54.9 (93.0%)
β‘ 26Γ faster query encoding on a single CPU thread (87 ms vs 2.3 s)
πΎ Matches score distillation while reading 12.6Γ less cached teacher data
π Paper: ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport (2609.34899)
π€ Checkpoints (all five students):
π» Code: https://github.com/Ryenhails/NanoVDR
π§© Single-vector predecessor, NanoVDR: https://arxiv.org/abs/2603.12824
If you already serve one of these teachers, swap in the matching student and keep your index as is. Feedback and upvotes welcome! π