Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches the teacher's MaxSim scores and so requires encoding and caching every training page, which can reach terabytes of page tokens. NanoVDR avoids pages entirely by training on the teacher's query embeddings alone, but only for single-vector retrievers. We present ColNanoVDR, to our knowledge the first framework to bring this document-free distillation to multi-vector VDR. Its objective, OTW (Optimal Transport with Learned Weights), aligns the student's query tokens with the teacher's by entropic optimal transport, with a learned weight for each student token, and needs no correspondence between the two tokenizations. We prove that the resulting alignment cost bounds the MaxSim score difference on every page. Distilled from five state-of-the-art teachers, the 149M text-only students retain about 95% of their teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster. Under identical training, OTW matches score distillation while encoding no page and reading 12.6x less cached teacher data.
Figures & tables
Figure 1: ColNanoVDR approaches its teacher’s quality at a fraction of its inference and training cost. (a) ViDoRe v3 NDCG@5 against query throughput on a single CPU thread. (b) Cached teacher data read during training: score distillation reads query and page tokens, OTW only query tokens.
Figure 2: Intuition behind OTW. The teacher and the student encode the same query into token sets of different sizes on the unit sphere. OTW aligns the two sets softly (shaded regions): each student token covers nearby teacher tokens, and its learned weight grows with the number it covers. On a page D never seen in training, aligned student and teacher tokens find the same best-matching page token, so the two MaxSim scores nearly coincide. Schematic.
Figure 3: The ColNanoVDR pipeline. Training (top): the student outputs unit-norm query tokens {si} and token weights a(θ) ; OTW (dashed box) aligns them with the cached teacher query tokens {tj} by entropic optimal transport, and the loss is the soft alignment cost ⟨Pε,C⟩ . Inference (bottom): the teacher is discarded; weighted student tokens aisi are scored with standard MaxSim against the unchanged teacher index.
Model
Params
v1
v2
v3
Reference systems (native retrieval)
Tomoro-ColQwen3-8B
8.8B
90.6
65.0
59.0
ColVec1.1-8b
8.4B
91.5
67.8
62.6
ColNomic-7B
7.8B
89.8
60.4
55.9
ColQwen3.5-4.5B
4.5B
91.6
63.7
58.7
Vultron-4.5B
4.5B
91.8
67.6
61.0
Table 1: Main results. NDCG@5 per benchmark and, for ColNanoVDR, retention of its own teacher in parentheses, each student scored against that teacher’s index. Reference systems are evaluated by us under the identical protocol (model identifiers in Appendix B.5 ).
Encode (ms)
Query encoder
Params
CPU
GPU
v3
Vision-language retrievers
Tomoro-ColQwen3-8B
8.8B
4,277
18.4
59.0
ColNomic-7B
7.8B
3,845
23.4
55.9
ColQwen3.5-4.5B (teacher)
4.5B
2,290
110.7 †
58.7
Tomoro-ColQwen3-4B
4.4B
2,118
18.6
57.6
Table 2: Query-encoding cost on one node, median over 20 queries at batch size 1, excluding MaxSim scoring: one CPU thread (float32) and one H200 (bf16); v3 is ViDoRe v3 NDCG@5. ColNanoVDR students of the same size share the encoder, so one row per size; v3 is that of the ColQwen3.5 student. † Torch fallback for the hybrid linear-attention layers. Protocol in Appendix C.1 .
Figure 4: ViDoRe v3 retention across the Ettin family, for the ColQwen3.5 and the Tomoro-ColQwen3-8B students.
ColQwen3.5-4.5B teacher
Tomoro-ColQwen3-8B teacher
Objective
Reads
Weights
v1
v2
v3
v1
v2
v3
document-dependent
InfoNCE
D
uniform
88.9
51.7
47.1
87.9
52.1
46.2
Listwise KL
Q+D
uniform
90.9
61.2
55.0
89.0
57.3
49.9
document-free
Coverage
Q
uniform
90.2
58.3
53.9
88.5
57.0
53.0
Table 3: Training objectives on two teachers (NDCG@5; same 149M Ettin student and training setup throughout). “Reads”: cached teacher data per step, query tokens (Q), document tokens with in-batch negatives (D), or both (Q+D). Bold: best per column.
Pool factor f
Vec./page
Teacher
Student
Ret.
1 (uncompressed)
1,764
61.6
59.1
95.8
3
587
61.6
59.0
95.7
9
195
61.1
58.2
95.3
Table 4: Index compression on ViDoRe v3 (mean NDCG@5 over its 8 datasets; retention in %). The ColVec1.1-4b index is pooled with hierarchical token pooling at pool factor f ; the teacher and its 149M student are scored against the same pooled index.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: The setting of Appendix A : the teacher’s query tokens form a uniform measure μT on the sphere, the student’s a weighted measure μS , and a transport plan aligns them. By Proposition A.2 , the cost of any such plan bounds the difference between the two weighted MaxSim scores on every page.
Part
Items
Token vectors
Size
queries, base
711,603
20,388,709
12.2 GiB
queries, translated
777,649
25,398,098
15.1 GiB
page images
711,603
529,150,162
315.4 GiB
Appendix
Table 5: ColQwen3.5 teacher cache written to disk for training (float16, 320-dimensional tokens). The translated variants add queries only.
Student
v1
v2
v3
ColQwen3.5-4.5B teacher
Ettin-32M
89.8 (98.0)
55.2 (86.7)
51.0 (87.0)
Ettin-68M
90.4 (98.7)
59.2 (92.9)
54.0 (92.0)
Ettin-150M
90.7 (99.0)
60.0 (94.2)
55.1 (93.8)
Ettin-400M
91.2 (99.5)
62.1 (97.4)
56.7 (96.5)
Tomoro-ColQwen3-8B teacher
Appendix
Table 6: Retention across the Ettin family for both ablation teachers (NDCG@5; % retention of that teacher).
Uncompressed ( f=1 )
f=3
f=9
Dataset
Pages
Teacher
Student
Ret.
Teacher
Student
Ret.
Teacher
Student
Ret.
Finance (en)
2,942
66.9
63.3
94.6
66.5
62.8
94.5
66.3
62.8
94.8
Finance (fr)
2,384
48.9
47.4
96.8
48.8
47.2
96.8
48.2
46.1
95.8
Computer sci.
1,360
78.1
74.2
95.1
77.9
74.2
95.2
77.6
73.5
94.8
Human res.
1,110
64.8
61.6
95.1
64.7
61.6
95.2
64.5
61.0
94.6
Energy
2,225
66.4
65.2
98.2
66.6
64.8
97.4
65.9
63.8
96.9
Appendix
Table 7: Index compression on ViDoRe v3 (NDCG@5; retention in %). The ColVec1.1-4b index is pooled with hierarchical token pooling at pool factor f , and the teacher and its 149M student are scored against the same pooled index. Average vectors per page: 1,764 ( f=1 ), 587 ( f=3 ), 195 ( f=9 ).
Head target
v1
v2
v3
hard assignment
90.7 (99.0)
60.0 (94.1)
55.1 (93.9)
soft assignment
90.7 (99.0)
60.1 (94.3)
55.1 (93.9)
OTW, joint
90.7 (99.0)
60.0 (94.2)
55.1 (93.8)
Appendix
Table 8: Two-stage weight heads on frozen coverage geometry, 149M student (NDCG@5; % retention). The OTW row is the jointly trained reference.
149M student
395M student
Weights
v1
v2
v3
v1
v2
v3
Learned (default)
90.7
60.0
55.1
91.2
62.1
56.7
Dead tokens pruned
90.7
60.0
54.9
91.1
62.2
56.4
Uniform
89.1
56.5
51.6
89.8
60.8
54.8
Inverted (dead only)
73.8
27.0
26.2
80.0
34.3
31.2
Appendix
Table 9: Inference-time weight interventions (NDCG@5): embeddings fixed, only the per-token weights replaced; dead tokens as defined in Appendix D.3 .
Weights
149M (coverage)
395M
hard assignment
55.2
56.6
soft, ε=0.005
55.2
56.6
soft, ε=0.02
55.2
56.7
soft, ε=0.05
55.3
56.7
soft, ε=0.1
55.3
56.6
soft, ε=0.2
55.0
56.3
Appendix
Table 10: Inference-weight variants on fixed geometries, NDCG@5 on ViDoRe v3. “149M (coverage)” is the 149M student trained with coverage; “395M” is the 395M OTW student. All variants except learned and uniform use the teacher’s query tokens at inference.
Objective
Reads
supD∣ΔSˉ∣
centered
ρs
W1
2OTc
2⟨Pε,C⟩
InfoNCE
D
0.335
0.251
0.780
1.256
1.262
1.287
Listwise KL
Q+D
0.157
0.106
0.936
0.913
0.925
0.953
Coverage
Q
0.084
0.079
0.949
0.613
0.656
0.694
OT-uniform
Q
0.083
0.082
0.945
0.656
0.674
0.723
OTW
Q
0.069
0.059
0.956
0.568
0.596
0.621
Appendix
Table 11: Measured discrepancy against the bounds of Corollary 2 for the ColQwen3.5 students of Table 3 , medians over the analysis sample. “centered” removes the per-query mean offset; ρs is the Spearman correlation between the student’s and the teacher’s document scores; “Reads” as in Table 3 .