Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.
Figures & tables
Figure 1: A page inverted from its stored index. The vector store keeps only the page’s index, one 320-dimensional vector per image patch. Inverted from that index, the page is recovered with its tables, illustrations and text. The page (from the ViDoRe v3 industrial subset) was chosen to show a mixed layout; its word recall is 0.620.
Figure 2: The three steps of the attack. Flames: models the attacker trains for an encoder E ; snowflakes: frozen models, which at inference include all of them. Step 0 identifies E from the stored index with public encoders alone (§ 4.2 ); step 1 trains models for E on public pages; in step 2, each row shows, per storage configuration, where the page shape and the order come from. The inverter uses its ordered variant when the order is known and its set variant otherwise (§ 4.1 ).
Index
Word rec. ↑
NED ↓
Sens. rec. ↑
Re-identification among 19,252
top-1
top-5
MRR
med. rank
indep.
Null model
–
0.001
0.958
0.008
0.000
0.000
0.001
8,161
–
kNN baseline
raw
0.234
0.858
0.146
0.199
0.438
0.320
7
–
Ours
raw
0.474
0.292
0.450
0.984
0.995
0.989
1
0.977
Ours
pooled × 3
0.076
0.858
0.044
0.011
0.035
0.029
496
0.005
Ours
pooled × 9
0.081
0.879
0.054
0.005
0.024
0.018
724
0.003
Table 1: Inverting the index of 2,000 held-out ViDoRe v3 pages (250 per domain) under EA (OCR columns over the 1,927 with ≥20 reference words). The raw, pooled × 3 and × 9 indices hold on average 1,241, 413 and 137 vectors per page and keep 100%, 99.2% and 97.8% of nDCG@10. Shapes are inferred from the vectors (§ 6.1 ), except for the shuffled row without the position model, which is given the true shape. Indep.: top-1 under the independent judge EB ; for the VAE row, of the source pages themselves. Null row: 200-page subset (Appendix I ). Bootstrap intervals: Appendix J ; precision and F1: Table 3 .
Figure 3: Numeric cells recovered from a raw index. A band of a statistical table on a ViDoRe v3 HR page, cut from the source page and from the page inverted from its raw index: 59 of its 60 numeric cells are recovered exactly. Yellow: tokens that sensitive-token recall counts; vermillion: the two tokens that differ, one numeric cell and one country code. The page was chosen by inspection and, with the two bands of Appendix A , gives one page per class of sensitive token; the band is its densest in sensitive tokens.
Figure 4: Our attacks do not invert pooled indices; a shuffled order can be restored. Left: one page under the three storage configurations (shuffled: without and with the position model), with its word recall and rank among 19,252; the page is typical by a fixed rule, its word recall closest to the mean under both the raw and the restored index (Appendix A ). Right: all 2,000 pages; dashed, the kNN baseline.
Index ( EB )
Word rec. ↑
NED ↓
Sens. rec. ↑
top-1 ↑
med. rank ↓
Null model
0.001
0.963
0.006
0.000
8,892.5
kNN baseline
0.223
0.860
0.114
0.154
12
raw, inferred shape
0.079
0.718
0.085
0.702
1
raw, true shape
0.094
0.655
0.098
0.818
1
shuffled
w/o position model
0.028
0.846
0.041
0.006
1,047
Table 2: The same attack and 2,000 pages (250 per domain) under EB . The shape is inferred (0.795 correct for the raw index), except for the shuffled row without the position model, which is given the true shape; the true-shape row isolates the cost of inference. Under the independent judge EA , top-1 is 0.605 (raw) and 0.008 and 0.076 (shuffled, w/o and w/ position model). EA row: a control inverter for EA trained on the same data as the EB inverter (top-1 0.980 under the judge EB ). Null row: 200-page subset.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Two further bands, chosen with Figure 3 as one page per class of sensitive token: prose with acronyms ( pharma ) and a company filing with its registration numbers ( finance_fr ). Source left, inverted page from the raw index right, marked as in Figure 3 ; a word is boxed only where OCR and the eye agree that it differs, so the registration number, which OCR misreads, is not. Both pages are ranked first among 19,252.
Figure 6: Whole pages, computer science. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 7: Whole pages, pharmaceuticals. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 8: Whole pages, physics. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 9: Whole pages, industrial. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 10: Whole pages, energy. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 11: Whole pages, human resources. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 12: Whole pages, finance (en). The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 13: Whole pages, finance (fr). The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
words
sensitive tokens
encoder
row
P
R
F1
P
R
F1
EA
Null model (200 pages)
0.008
0.001
0.001
0.090
0.008
0.010
kNN baseline
0.220
0.234
0.213
0.131
0.146
0.128
Ours, raw
0.475
0.474
0.474
0.455
0.450
0.450
Ours, pooled × 3
0.078
0.076
0.074
0.053
0.044
0.044
Ours, pooled × 9
0.076
0.081
0.076
0.056
0.054
0.051
Appendix
Table 3: Precision, recall and F1 of the recovered words and sensitive tokens, as means over pages, on the same pages and transcripts as Tables 1 and 2 . Precision is the share of the inverted page’s words (or sensitive tokens) that occur in the source page’s transcript, counted as multisets. The null rows and the VAE rows use the 200-page subset.
pages per decision
1
10
100
Pool of 13: candidates kept, top-1 accuracy
structure, candidates kept
2.77
2.77
2.77
query probe
0.999
1.000
1.000
reference cloud
1.000
1.000
1.000
True encoder removed: query probe
flagged as outside the pool
0.863
1.000
–
Appendix
Table 4: Identifying the encoder of each of the 13 target stores from the indices of 1, 10 or 100 of its pages (26,000, 2,600 and 260 decisions). Structure keeps the true encoder in every decision. The lower block removes each store’s encoder from the pool in turn; from 100 pages the threshold set on development stores rejects most in-pool decisions on the indexed pages, so no rate is given.
encoder
demb
vectors per page
vector count
TomoroAI/tomoro-colqwen3-embed-8b
EA
320
1,240 (936–1,276)
pixel budget, ≤ 1,280
TomoroAI/tomoro-colqwen3-embed-4b
320
1,240 (936–1,276)
pixel budget, ≤ 1,280
athrael-soju/colqwen3.5-4.5B-v3
EB
320
744 (720–768)
pixel budget, ≤ 768
vultr/VultronRetrieverCore-Qwen3.5-4.5B
320
1,240 (936–1,276)
pixel budget, ≤ 1,280
webAI-Official/webAI-ColVec1.1-4b
640
1,750 (936–1,776)
pixel budget, ≤ 1,792
webAI-Official/webAI-ColVec1.1-8b
640
1,750 (936–1,776)
pixel budget, ≤ 1,792
Appendix
Table 5: The encoder pool. Vectors per page: median and range over the fixed evaluation set, image tokens only where they form one span; for the two ColSmol models and ColLFM2 they do not, and the whole output is the index. Vector count: the counts the encoder’s processor can emit, which the structure test checks.
Figure 14: The latent flow-matching inverter : a double-stream MMDiT in which the stored index takes the place of the text stream, with separate weights per stream and joint attention in each of the 16 blocks. Solid: a training step on Eq. ( 1 ); dashed: sampling, 50 Euler steps from noise and then the frozen VAE decoder. The condition stream carries 2D rotary positions for the raw index and for the shuffled index after the position model, and none for the pooled index and the shuffled index without the position model.
Figure 15: Every trained model and the results it produces. Each row is a stored index the attacker meets; it gives how the page shape and the order are obtained, which inverter reads the index, and where the result is reported. The rows with the position model reuse the raw ordered inverter (A1, B1) unchanged; the rows without it are given the true page shape. Training details are in Table 6 .
model
reads
training pages
parameters
steps × batch
hardware
time
A1, ordered inverter
raw index of EA
682,818
337.1M
100k × 16
1 H200
23.3 h
A2, set inverter
pooled × 3 index of EA
682,818
337.1M
100k × 16
1 H200
19.2 h
A3, set inverter
pooled × 9 index of EA
682,818
337.1M
100k × 16
1 H200
18.0 h
A4, set inverter
shuffled index of EA
682,818
337.1M
100k × 16
1 H200
22.9 h
B1, ordered inverter
raw index of EB
346,770
337.1M
100k × 16
1 H200
12.9 h
B2, set inverter
shuffled index of EB
346,770
337.1M
100k × 16
1 H200
12.7 h
Appendix
Table 6: Training setup of every model the attacker trains. A5 is the control of § 6.4 , trained on the VisRAG part alone (Appendix B ). Inverters: Eq. ( 1 ), condition dropout 0.1, AdamW ( β=(0.9,0.95) , weight decay 0.01) at 10−4 with 1k warmup and cosine decay to 10%, gradient clipping at 1.0, bf16; validation loss on 640 pages spaced evenly through the in-distribution validation split. Set classifiers: cross-entropy over 31 page shapes with logits masked to the shapes compatible with the vector count, AdamW (weight decay 0.01) at 10−3 with one-cycle scheduling. Position models: squared error on each vector’s normalised row and column, and on the log aspect ratio for the 30% of pages whose shape is withheld, AdamW (weight decay 0.01) at 3×10−4 with 1k warmup and cosine decay, bf16; the 0.43M per-vector probe is trained jointly on the same batches, and both are validated on 1,024 validation pages. Times are wall-clock times of the training jobs.
Figure 16: Word recall and re-identification under position noise. Word recall (top) and top-1 (bottom) of the unchanged ordered inverter against the mean position error of the index it is given, under EA . Grey: the true positions with Gaussian noise ( σ=0.5,1,2,4 ); green: the position models on the shuffled index, the position model (P2) filled and the per-vector probe (P1) open; dashed: the true order, the same shuffled index without a position model, and a random order.
Figure 17: Validation loss of set and ordered inverters. Validation loss of four inverters trained identically: on the ordered raw index, on the same vectors with their order shuffled, and on the pooled × 3 and × 9 indices. The three unordered variants coincide throughout training; only the ordered index is learnable by this inverter.
Control (200 pages)
EA
EB
× 9
null, word recall
0.001
0.001
0.001
null, NED
0.958
0.963
0.943
seed agreement (F1)
0.476
0.130
0.094
vs. truth (F1)
0.480
0.098
0.076
recall, true shape
0.481
0.094
0.083
recall, inferred shape
0.478
0.080
0.083
Appendix
Table 7: Controls separating content determined by the index from content generated by the prior, on a 200-page subset spaced evenly through the fixed set. EA and EB are the raw index of each encoder; × 9 is the pooled index of EA . All three columns are measured at the reported guidance scale. The null rows are scored with the Latin-script recogniser, as in the main tables; the seed-agreement and recall rows come from the sampling run and its English recogniser, which changes text metrics by at most 0.01 (Appendix C ). The VAE reconstruction is the last row of Table 1 ; near-duplicates between the training and target corpora are quantified in § 5 .
Index, attack
word recall
NED
sens. recall
top-1
kNN (nearest train page)
0.234 [0.227, 0.241]
0.858 [0.853, 0.864]
0.146 [0.140, 0.152]
0.199 [0.181, 0.216]
raw, ordered
0.474 [0.468, 0.480]
0.292 [0.285, 0.299]
0.450 [0.442, 0.457]
0.984 [0.978, 0.990]
raw, ordered, true shape
0.476 [0.470, 0.483]
0.290 [0.283, 0.296]
0.450 [0.443, 0.457]
0.986 [0.981, 0.991]
shuffled, w/o position model
0.084 [0.081, 0.087]
0.842 [0.838, 0.846]
0.048 [0.045, 0.050]
0.038 [0.030, 0.047]
shuffled, w/ position model
0.231 [0.226, 0.236]
0.925 [0.921, 0.929]
0.212 [0.206, 0.217]
0.935 [0.924, 0.946]
pooled × 3
0.076 [0.073, 0.079]
0.858 [0.853, 0.862]
0.044 [0.042, 0.046]
0.011 [0.006, 0.015]
Appendix
Table 8: Means over pages with 95% bootstrap intervals under EA (top), and paired differences on the same pages (bottom). The null row of Table 1 is a 200-page control and is not resampled.
Figure 18: Re-identification versus word recall. The 1,927 scored pages of the raw index under EA , binned by word recall: the share whose source ranks first, second to tenth, or beyond tenth among 19,252, with the number of pages per bin above each bar. Below a word recall of 0.1, 75% of pages still rank first; over all pages the Spearman correlation between the two is ρ=0.13 .
Subset
top-1
top-5
MRR
cs (textbooks)
1.000
1.000
1.000
energy
0.996
0.996
0.997
finance_en
0.996
0.996
0.997
physics (slides)
0.992
1.000
0.995
hr
0.984
0.992
0.987
pharma
0.980
1.000
0.989
Appendix
Table 9: Re-identification of raw-index inversions under EA by ViDoRe v3 subset (250 pages each, 19,252-page corpus). Judged by the independent encoder EB instead: top-1 0.977.
Figure 19: One page of the fixed set (ViDoRe v3 cs ): source (left), the ordered inverter on the shuffled index after the position model restores the order (middle), and the same inverter on the true order (right), each with its word recall, NED and rank. Same inverter, seed and guidance scale.
Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings per page incurs substantial overhead. Existing training-free methods rely on pruning or merging: pruning degrades sharply under aggressive compression, whereas merging does not explicitly prioritize important regions when forming representatives. We introduce AnchorFold, a training-free focus-then-fold framework for document-side index compression. AnchorFold applies Recursive Attention Propagation over visual self-attention graphs, performing multi-step propagation within each attention head and integrating scores across heads and layers. The focus stage selects the highest-centrality tokens as anchors. The fold stage assigns remaining tokens to their most similar anchors in the normalized retrieval space and summarizes each anchor-centered group through centrality-weighted aggregation. This preserves non-anchor contributions while concentrating capacity on structurally important tokens. Across ViDoRe v1/v2 and REAL-MM-RAG with three diverse retrieval backbones, AnchorFold consistently outperforms all evaluated training-free baselines at γ≤0.20. On ViDoRe v1/v2, it retains 98.3% of full-index NDCG@5 on average at 5× compression, achieving near-lossless compression, and 92.4% at 20× compression.
Haoyu Zuo, Yibo Yan, Xin Zou +4
Hong Kong University of Science and Technology (Guangzhou) · Alibaba Cloud Computing · Hong Kong University of Science and Technology
Visual document retrieval has become essential for accessing information in visually rich documents. Existing approaches fall into two camps. Late-interaction retrievers achieve strong quality through fine-grained token-level matching but store hundreds of vectors per page, incurring large index footprints and high serving costs. By contrast, dense single-vector retrievers retain storage and latency advantages but consistently lag in quality because they compress all information into a single final-layer embedding. In this work, we first conduct a layerwise diagnostic on single-vector retrievers, revealing that retrieval-relevant signal resides in internal representations. Motivated by these findings, we propose MINER (Mining Multimodal Internal RepreseNtation for Efficient Retrieval), a lightweight plug-in module that probes and fuses internal signals across transformer layers into a single compact embedding without modifying the backbone or sacrificing single-vector efficiency. The first Retrieval-Aligned Layer Probing stage attaches a lightweight probe at each layer, surfacing which dimensions carry retrieval-relevant information. The subsequent Adaptive Sparse Multi-Layer Fusion stage applies performance-adaptive neuron-level masking to the selected layers and fuses the surviving signals into the final dense vector. Across ViDoRe V1/V2/V3, MINER outperforms existing dense single-vector retrievers on the majority of benchmarks, with up to 4.5% nDCG@5 improvement over its corresponding backbone. Compared to strong late-interaction baselines, in some settings MINER substantially narrows the nDCG@5 gap to 0.2 while preserving the storage and serving advantages of dense retrieval.
Weien Li, Rui Song, Zeyu Li +8
McGill University · MIT - Massachusetts Institute of Technology · University of Cambridge +3
Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either train a smaller multi-vector encoder from scratch or distil only the query side; neither yields a compact single-vector retriever end-to-end. We present DistilVDR, a 524M end-to-end VDR system distilled bilaterally from a single 8B vision-language teacher under a pointwise cosine alignment loss. All supervision comes from the frozen teacher's embedding space, which was itself trained with relevance supervision, so the student objective needs no relevance labels, negative sampling, or contrastive term. We match VDR's text-query and image-document input asymmetry with an asymmetric encoder-only student that concentrates visual capacity on the document side and keeps the query side at 70M parameters. We release two variants that share the same encoders and training and differ only in the document encoder's visual-tile budget: DistilVDR-HiRes attains 61.74 average NDCG@5 on ViDoRe v1+v2+v3 (86.9% of the 8B teacher) and leads every reproduced sub-1B baseline on the high-resolution-sensitive v3 benchmark, while DistilVDR-Fast attains 59.98 at a 3 times smaller visual-token budget. Both variants store one million documents in a 15.6 times smaller index than the strongest sub-1B multi-vector baseline and index the corpus an order of magnitude faster. The code is available at https://github.com/Ryenhails/NanoVDR.
Zhuchenyang Liu, Ziyi Wang, Yao Zhang +1
Aalto University, Finland · Independent Researcher, Netherlands