Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.
Figures & tables
Figure 1: A page inverted from its stored index. The vector store keeps only the page’s index, one 320-dimensional vector per image patch. Inverted from that index, the page is recovered with its tables, illustrations and text. The page (from the ViDoRe v3 industrial subset) was chosen to show a mixed layout; its word recall is 0.620.
Figure 2: The three steps of the attack. Flames: models the attacker trains for an encoder E ; snowflakes: frozen models, which at inference include all of them. Step 0 identifies E from the stored index with public encoders alone (§ 4.2 ); step 1 trains models for E on public pages; in step 2, each row shows, per storage configuration, where the page shape and the order come from. The inverter uses its ordered variant when the order is known and its set variant otherwise (§ 4.1 ).
Index
Word rec. ↑
NED ↓
Sens. rec. ↑
Re-identification among 19,252
top-1
top-5
MRR
med. rank
indep.
Null model
–
0.001
0.958
0.008
0.000
0.000
0.001
8,161
–
kNN baseline
raw
0.234
0.858
0.146
0.199
0.438
0.320
7
–
Ours
raw
0.474
0.292
0.450
0.984
0.995
0.989
1
0.977
Ours
pooled × 3
0.076
0.858
0.044
0.011
0.035
0.029
496
0.005
Ours
pooled × 9
0.081
0.879
0.054
0.005
0.024
0.018
724
0.003
Table 1: Inverting the index of 2,000 held-out ViDoRe v3 pages (250 per domain) under EA (OCR columns over the 1,927 with ≥20 reference words). The raw, pooled × 3 and × 9 indices hold on average 1,241, 413 and 137 vectors per page and keep 100%, 99.2% and 97.8% of nDCG@10. Shapes are inferred from the vectors (§ 6.1 ), except for the shuffled row without the position model, which is given the true shape. Indep.: top-1 under the independent judge EB ; for the VAE row, of the source pages themselves. Null row: 200-page subset (Appendix I ). Bootstrap intervals: Appendix J ; precision and F1: Table 3 .
Figure 3: Numeric cells recovered from a raw index. A band of a statistical table on a ViDoRe v3 HR page, cut from the source page and from the page inverted from its raw index: 59 of its 60 numeric cells are recovered exactly. Yellow: tokens that sensitive-token recall counts; vermillion: the two tokens that differ, one numeric cell and one country code. The page was chosen by inspection and, with the two bands of Appendix A , gives one page per class of sensitive token; the band is its densest in sensitive tokens.
Figure 4: Our attacks do not invert pooled indices; a shuffled order can be restored. Left: one page under the three storage configurations (shuffled: without and with the position model), with its word recall and rank among 19,252; the page is typical by a fixed rule, its word recall closest to the mean under both the raw and the restored index (Appendix A ). Right: all 2,000 pages; dashed, the kNN baseline.
Index ( EB )
Word rec. ↑
NED ↓
Sens. rec. ↑
top-1 ↑
med. rank ↓
Null model
0.001
0.963
0.006
0.000
8,892.5
kNN baseline
0.223
0.860
0.114
0.154
12
raw, inferred shape
0.079
0.718
0.085
0.702
1
raw, true shape
0.094
0.655
0.098
0.818
1
shuffled
w/o position model
0.028
0.846
0.041
0.006
1,047
Table 2: The same attack and 2,000 pages (250 per domain) under EB . The shape is inferred (0.795 correct for the raw index), except for the shuffled row without the position model, which is given the true shape; the true-shape row isolates the cost of inference. Under the independent judge EA , top-1 is 0.605 (raw) and 0.008 and 0.076 (shuffled, w/o and w/ position model). EA row: a control inverter for EA trained on the same data as the EB inverter (top-1 0.980 under the judge EB ). Null row: 200-page subset.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Two further bands, chosen with Figure 3 as one page per class of sensitive token: prose with acronyms ( pharma ) and a company filing with its registration numbers ( finance_fr ). Source left, inverted page from the raw index right, marked as in Figure 3 ; a word is boxed only where OCR and the eye agree that it differs, so the registration number, which OCR misreads, is not. Both pages are ranked first among 19,252.
Figure 6: Whole pages, computer science. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 7: Whole pages, pharmaceuticals. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 8: Whole pages, physics. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 9: Whole pages, industrial. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 10: Whole pages, energy. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 11: Whole pages, human resources. The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 12: Whole pages, finance (en). The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
Figure 13: Whole pages, finance (fr). The pages of this domain at the 75th, 50th and 25th percentile of word recall under the raw index, and the pages inverted from the raw, the pooled × 9 and the shuffled index with the position model; each inverted page carries its own word recall and rank among 19,252.
words
sensitive tokens
encoder
row
P
R
F1
P
R
F1
EA
Null model (200 pages)
0.008
0.001
0.001
0.090
0.008
0.010
kNN baseline
0.220
0.234
0.213
0.131
0.146
0.128
Ours, raw
0.475
0.474
0.474
0.455
0.450
0.450
Ours, pooled × 3
0.078
0.076
0.074
0.053
0.044
0.044
Ours, pooled × 9
0.076
0.081
0.076
0.056
0.054
0.051
Appendix
Table 3: Precision, recall and F1 of the recovered words and sensitive tokens, as means over pages, on the same pages and transcripts as Tables 1 and 2 . Precision is the share of the inverted page’s words (or sensitive tokens) that occur in the source page’s transcript, counted as multisets. The null rows and the VAE rows use the 200-page subset.
pages per decision
1
10
100
Pool of 13: candidates kept, top-1 accuracy
structure, candidates kept
2.77
2.77
2.77
query probe
0.999
1.000
1.000
reference cloud
1.000
1.000
1.000
True encoder removed: query probe
flagged as outside the pool
0.863
1.000
–
Appendix
Table 4: Identifying the encoder of each of the 13 target stores from the indices of 1, 10 or 100 of its pages (26,000, 2,600 and 260 decisions). Structure keeps the true encoder in every decision. The lower block removes each store’s encoder from the pool in turn; from 100 pages the threshold set on development stores rejects most in-pool decisions on the indexed pages, so no rate is given.
encoder
demb
vectors per page
vector count
TomoroAI/tomoro-colqwen3-embed-8b
EA
320
1,240 (936–1,276)
pixel budget, ≤ 1,280
TomoroAI/tomoro-colqwen3-embed-4b
320
1,240 (936–1,276)
pixel budget, ≤ 1,280
athrael-soju/colqwen3.5-4.5B-v3
EB
320
744 (720–768)
pixel budget, ≤ 768
vultr/VultronRetrieverCore-Qwen3.5-4.5B
320
1,240 (936–1,276)
pixel budget, ≤ 1,280
webAI-Official/webAI-ColVec1.1-4b
640
1,750 (936–1,776)
pixel budget, ≤ 1,792
webAI-Official/webAI-ColVec1.1-8b
640
1,750 (936–1,776)
pixel budget, ≤ 1,792
Appendix
Table 5: The encoder pool. Vectors per page: median and range over the fixed evaluation set, image tokens only where they form one span; for the two ColSmol models and ColLFM2 they do not, and the whole output is the index. Vector count: the counts the encoder’s processor can emit, which the structure test checks.
Figure 14: The latent flow-matching inverter : a double-stream MMDiT in which the stored index takes the place of the text stream, with separate weights per stream and joint attention in each of the 16 blocks. Solid: a training step on Eq. ( 1 ); dashed: sampling, 50 Euler steps from noise and then the frozen VAE decoder. The condition stream carries 2D rotary positions for the raw index and for the shuffled index after the position model, and none for the pooled index and the shuffled index without the position model.
Figure 15: Every trained model and the results it produces. Each row is a stored index the attacker meets; it gives how the page shape and the order are obtained, which inverter reads the index, and where the result is reported. The rows with the position model reuse the raw ordered inverter (A1, B1) unchanged; the rows without it are given the true page shape. Training details are in Table 6 .
model
reads
training pages
parameters
steps × batch
hardware
time
A1, ordered inverter
raw index of EA
682,818
337.1M
100k × 16
1 H200
23.3 h
A2, set inverter
pooled × 3 index of EA
682,818
337.1M
100k × 16
1 H200
19.2 h
A3, set inverter
pooled × 9 index of EA
682,818
337.1M
100k × 16
1 H200
18.0 h
A4, set inverter
shuffled index of EA
682,818
337.1M
100k × 16
1 H200
22.9 h
B1, ordered inverter
raw index of EB
346,770
337.1M
100k × 16
1 H200
12.9 h
B2, set inverter
shuffled index of EB
346,770
337.1M
100k × 16
1 H200
12.7 h
Appendix
Table 6: Training setup of every model the attacker trains. A5 is the control of § 6.4 , trained on the VisRAG part alone (Appendix B ). Inverters: Eq. ( 1 ), condition dropout 0.1, AdamW ( β=(0.9,0.95) , weight decay 0.01) at 10−4 with 1k warmup and cosine decay to 10%, gradient clipping at 1.0, bf16; validation loss on 640 pages spaced evenly through the in-distribution validation split. Set classifiers: cross-entropy over 31 page shapes with logits masked to the shapes compatible with the vector count, AdamW (weight decay 0.01) at 10−3 with one-cycle scheduling. Position models: squared error on each vector’s normalised row and column, and on the log aspect ratio for the 30% of pages whose shape is withheld, AdamW (weight decay 0.01) at 3×10−4 with 1k warmup and cosine decay, bf16; the 0.43M per-vector probe is trained jointly on the same batches, and both are validated on 1,024 validation pages. Times are wall-clock times of the training jobs.
Figure 16: Word recall and re-identification under position noise. Word recall (top) and top-1 (bottom) of the unchanged ordered inverter against the mean position error of the index it is given, under EA . Grey: the true positions with Gaussian noise ( σ=0.5,1,2,4 ); green: the position models on the shuffled index, the position model (P2) filled and the per-vector probe (P1) open; dashed: the true order, the same shuffled index without a position model, and a random order.
Figure 17: Validation loss of set and ordered inverters. Validation loss of four inverters trained identically: on the ordered raw index, on the same vectors with their order shuffled, and on the pooled × 3 and × 9 indices. The three unordered variants coincide throughout training; only the ordered index is learnable by this inverter.
Control (200 pages)
EA
EB
× 9
null, word recall
0.001
0.001
0.001
null, NED
0.958
0.963
0.943
seed agreement (F1)
0.476
0.130
0.094
vs. truth (F1)
0.480
0.098
0.076
recall, true shape
0.481
0.094
0.083
recall, inferred shape
0.478
0.080
0.083
Appendix
Table 7: Controls separating content determined by the index from content generated by the prior, on a 200-page subset spaced evenly through the fixed set. EA and EB are the raw index of each encoder; × 9 is the pooled index of EA . All three columns are measured at the reported guidance scale. The null rows are scored with the Latin-script recogniser, as in the main tables; the seed-agreement and recall rows come from the sampling run and its English recogniser, which changes text metrics by at most 0.01 (Appendix C ). The VAE reconstruction is the last row of Table 1 ; near-duplicates between the training and target corpora are quantified in § 5 .
Index, attack
word recall
NED
sens. recall
top-1
kNN (nearest train page)
0.234 [0.227, 0.241]
0.858 [0.853, 0.864]
0.146 [0.140, 0.152]
0.199 [0.181, 0.216]
raw, ordered
0.474 [0.468, 0.480]
0.292 [0.285, 0.299]
0.450 [0.442, 0.457]
0.984 [0.978, 0.990]
raw, ordered, true shape
0.476 [0.470, 0.483]
0.290 [0.283, 0.296]
0.450 [0.443, 0.457]
0.986 [0.981, 0.991]
shuffled, w/o position model
0.084 [0.081, 0.087]
0.842 [0.838, 0.846]
0.048 [0.045, 0.050]
0.038 [0.030, 0.047]
shuffled, w/ position model
0.231 [0.226, 0.236]
0.925 [0.921, 0.929]
0.212 [0.206, 0.217]
0.935 [0.924, 0.946]
pooled × 3
0.076 [0.073, 0.079]
0.858 [0.853, 0.862]
0.044 [0.042, 0.046]
0.011 [0.006, 0.015]
Appendix
Table 8: Means over pages with 95% bootstrap intervals under EA (top), and paired differences on the same pages (bottom). The null row of Table 1 is a 200-page control and is not resampled.
Figure 18: Re-identification versus word recall. The 1,927 scored pages of the raw index under EA , binned by word recall: the share whose source ranks first, second to tenth, or beyond tenth among 19,252, with the number of pages per bin above each bar. Below a word recall of 0.1, 75% of pages still rank first; over all pages the Spearman correlation between the two is ρ=0.13 .
Subset
top-1
top-5
MRR
cs (textbooks)
1.000
1.000
1.000
energy
0.996
0.996
0.997
finance_en
0.996
0.996
0.997
physics (slides)
0.992
1.000
0.995
hr
0.984
0.992
0.987
pharma
0.980
1.000
0.989
Appendix
Table 9: Re-identification of raw-index inversions under EA by ViDoRe v3 subset (250 pages each, 19,252-page corpus). Judged by the independent encoder EB instead: top-1 0.977.
Figure 19: One page of the fixed set (ViDoRe v3 cs ): source (left), the ordered inverter on the shuffled index after the position model restores the order (middle), and the same inverter on the true order (right), each with its word recall, NED and rank. Same inverter, seed and guidance scale.