Visual instance retrieval often fails when the same object appears at different apparent sizes in the query and gallery. We show that the dominant cause is usually not resolution loss but object-to-image (O2I) ratio mismatch: the object occupies different fractions of the two images. On a controlled benchmark of 3,021 Objaverse objects rendered at five camera distances, more than 80% of the cross-distance degradation is attributable to O2I mismatch rather than resolution for 9 of 12 pretrained backbones; multi-scale architectures cut the resolution-only effect to single digits yet remain equally susceptible. The failure is also asymmetric: tight queries retrieve more reliably against wide gallery images than the reverse. Guided by this analysis, query-side scale augmentation and an OWLv2 crop reranker reach state of the art on ILIAS 100M (29.2 mAP@1000 before reranking, 42.0 after) without training or modifying the precomputed gallery index, and a LoRA fine-tune matches the query-side gains at a single forward pass, showing that O2I robustness is learnable.
Figures & tables
Figure 1: We study how object-to-image (O2I) ratio mismatch degrades feature similarity (panel A). Motivated by the analysis results, we apply query-side (panel B) and gallery-side (panel C) interventions to reduce the O2I mismatch between query and gallery, and observe a significant improvement in retrieval outcome, reaching state of the art on the ILIAS 100M benchmark.
Figure 2: A comparison of retrieval performance of twelve backbones as the O2I ratio changes. The y-axis is cross-distance R@1; the x-axis is the camera-distance ratio dq/dt (ratio =1 is self-retrieval). All backbones degrade with O2I mismatch, including those built for scale invariance (Swin-L, FlexiViT-L, MViTv2-L). Performance degrades for both directions, losing 0.5–0.9 R@1 at the extremes.
Figure 3: Disentangling O2I from resolution across backbones. X-axis: cross-distance Top-1 failure rate ( 1−R@1 ) on 3,021 Objaverse objects. We tight-crop both query and gallery to the object bbox to match O2I. The teal segment is the failure recovered by this matching. The orange segment remains, attributable to resolution alone. Trailing % is orange’s share of the total. For 9 of 12 backbones tested it is less than 20% : O2I mismatch dominates. Multi-scale backbones have small resolution shares yet remain equally susceptible to O2I. Per-backbone numbers: Table 2 ( Appendix C ).
Table 1: Simple query- and gallery-side object-to-image matching interventions achieve state of the art on ILIAS 100M. DINOv3-L backbone, 1,232 queries, ∼ 100M gallery. Oracle = upper bound for any reranker on the DINOv3-L top-1000 shortlist. Full table with query/gallery-side cost breakdown and paired t -test p -values: Table 13 ( Appendix N ).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Example renderings from the controlled benchmark. The same backpack rendered from the same viewpoint at five camera distances; each label shows the DINOv3-L cosine similarity to the reference image (green border, d=2.5 ). Cosine drops below 0.8 at every off-reference distance, despite the object and viewpoint being identical—changing only the object’s fraction of the frame is enough to break feature similarity.
Figure 5: Benchmark diversity. Top block: one instance from each of the 120 categories. Bottom two rows: instances of two example categories, illustrating within-category variety. All images are the reference-distance ( d=2.5 ) renders.
Backbone
Input px
α Raw cross-dist. R@1
β R@1 after O2I matching
γ=1−α Total drop
δ=1−β Resolution-only residual
ε=δ/γ Resolution share
Self-supervised
DINOv3-L
224
0.785
0.960
0.215
0.040
19%
DINOv2-L
224
0.752
0.952
0.248
0.048
19%
DINOv2-G
224
0.769
0.956
0.231
0.044
19%
Vision-language
SigLIP-2-L
384
0.798
0.887
0.202
0.113
56%
Appendix
Table 2: Disentangling O2I from resolution across backbones (3,021 Objaverse objects, mean cross-distance R@1 over 20 ordered distance pairs; both conditions evaluated against single-distance galleries of 3,021 items as described in Section 4.2 , so absolute values are higher than the multi-distance-gallery numbers of Table 5 ). α : standard retrieval (query and gallery differ in both object-to-image fraction, O2I, and pixel resolution). β : same retrieval after both query and gallery are tight-cropped to the object bbox and bilinearly upsampled to the backbone’s native input size, so only resolution differs. Derived columns ( γ , δ , ε ) are computed per the formulas in the header; the O2I share is 1−ε . For 9 of the 12 evaluated backbones (9 of 11 retrieval backbones), resolution accounts for ≤ 20% of the cross-distance failure. Hierarchical/multi-scale models (Swin-L, FlexiViT-L, MViTv2-L) have the smallest resolution-only residual, yet their O2I-only drops are comparable to standard ViTs of similar size—architectural scale invariance does not transfer to O2I invariance. DUSt3R-L is included as a reference (reconstruction, not retrieval); its failure is dominated by resolution loss, the opposite of the dominant retrieval pattern.
Cosine
c@256
c@512
c@1024
c@1536
c@256
1.00
0.96
0.87
0.77
c@512
0.96
1.00
0.94
0.83
c@1024
0.87
0.94
1.00
0.94
c@1536
0.77
0.83
0.94
1.00
Appendix
Table 3: O2I moves, resolution frozen. Mean cosine similarity between DINOv3-L features of the identical object crop pasted on gray canvases of different sizes (3,021 objects; bootstrap 95% CIs within ±0.003 ). Only the fraction of the image occupied by the object changes across canvases; the object’s pixels are byte-identical.
Figure 6: Cross-distance retrieval is asymmetric in O2I. Each axis reports mean MRR over the 10 ordered cross-distance pairs on our synthetic Objaverse benchmark. The x-axis is the direction in which the query is farther from the object than the gallery (MRR q>t ); the y-axis is the reverse direction (MRR q<t , query closer than the gallery). The dashed diagonal is y=x (no asymmetry). Each marker is one backbone; marker shape encodes the family, color shade the specific backbone within the family. Every backbone sits above the diagonal: a tight query against a wide gallery matches more reliably than the reverse.
Family
Backbone
MRR q<t
MRR q>t
Δ
tight query
wide query
Self-supervised
DINOv3-L
0.715
0.668
+0.047
DINOv2-L
0.689
0.614
+0.075
DINOv2-G
0.701
0.626
+0.075
Vision-language
SigLIP-2-L
0.719
0.699
+0.020
CLIP-L
0.726
0.686
+0.040
Appendix
Table 4: Cross-distance retrieval is asymmetric in O2I. Mean MRR per backbone over the ten ordered (dq,dt) cross-distance pairs in each direction. MRR q<t (“tight query”): query closer than target, query O2I > target O2I. MRR q>t (“wide query”): query farther than target, query O2I < target O2I. Δ = MRR q<t − MRR q>t is positive for every backbone: a tight query against a wide gallery matches more reliably than the reverse.
Table 5: O2I robustness: overall retrieval metrics. Mean over all 20 cross-distance pairs (3,021 objects, 15,101 gallery items). Best per column in bold .
Model
∣Δd∣ =0.5
∣Δd∣ =1.0
∣Δd∣ =1.5
∣Δd∣ =2.0
∣Δd∣ =2.5
∣Δd∣ =3.0
∣Δd∣ =3.5
Mean
Self-supervised
DINOv3-L
0.962
0.866
0.764
0.454
0.456
0.250
0.110
0.645
DINOv2-L
0.962
0.854
0.749
0.343
0.377
0.172
0.053
0.608
DINOv2-G
0.966
0.866
0.773
0.346
0.415
0.173
0.052
0.620
Vision-language
SigLIP-2-L
0.967
0.874
0.785
0.442
0.510
0.288
0.126
0.662
Appendix
Table 6: O2I robustness: R@1 by camera distance gap. Cross-distance retrieval over 3,021 objects at 5 camera distances. ∣Δd∣ is the absolute gap between query and target camera distances. The O2I mismatch ratio O2Iq/O2It=(dt/dq)2 is not monotone in ∣Δd∣ on this discrete grid (e.g., ∣Δd∣=2.5 corresponds to a 4× ratio, smaller than the 5.44× at ∣Δd∣=2.0 ); accordingly, R@1 at ∣Δd∣=2.5 exceeds R@1 at ∣Δd∣=2.0 for most backbones, consistent with the paper’s claim that O2I drives the failure. Gallery: target at dtarget + all other objects at all distances (15,101 items). Best per column in bold .
ROxford5K
RParis6K
Method
Easy
Med
Hard
Easy
Med
Hard
Original
93.5
81.5
64.2
95.3
91.6
82.2
GP (5 scales)
91.1
81.1
64.8
95.3
91.9
83.0
MSA
93.1
82.3
65.8
95.6
92.3
83.5
Appendix
Table 7: ROxford5K and RParis6K (DINOv3-L, 768 px). MSA improves mAP across both benchmarks, with the largest gains on the Hard protocol where O2I mismatch is most severe.
Method
R@1
R@5
R@10
mAP@1000
Original (baseline)
55.0
62.5
66.8
26.5
GP-5
57.9
66.6
69.6
28.3
MSA-3
59.2
68.3
70.5
29.2
Original + α QE
55.0
55.8
57.1
29.1
GP-5 + α QE
57.9
59.4
60.6
30.7
MSA-3 + α QE
59.2
60.8
62.1
31.4
Appendix
Table 8: Query-side O2I matching vs. query expansion on ILIAS 100M (DINOv3-L). α QE lifts mAP but cannot fix a wrong top match; O2I matching improves every metric, and the two stack.
Method
Views
R@1
R@5
mAP
Original
1
55.0
62.5
26.5
Horizontal flip
2
55.9
63.2
26.9
Color jitter
5
55.3
63.0
26.7
Random crops
5
55.5
62.7
26.4
Mixed non-scale
5
55.8
63.5
27.0
GP (scale)
5
58.2
65.9
28.3
Appendix
Table 9: Scale vs. generic augmentation on ILIAS 100M (DINOv3-L). GP provides 3× the R@1 gain of non-scale augmentations at equal query-time cost (5 views).
Method
R@1
R@5
R@10
mAP
Original (baseline)
55.0
62.5
66.8
26.5
Scale averaging variants
GP avg (5 scales)
58.2
65.9
69.1
28.3
Median across scales
57.4
65.4
68.8
28.2
Consistency-weighted mean
58.2
65.9
69.1
28.3
Object localization (requires bbox)
Appendix
Table 10: Ablation of query-side strategies on ILIAS 100M (DINOv3-L, 768 px gallery). All methods are training-free and query-only. MSA achieves the best results on all metrics.
mAP
Condition
O2I
Outside-bbox
INSTRE
ILIAS
Full image (baseline)
unchanged
attended
0.854
0.628
BBox attention mask
unchanged
masked
0.886
0.747
GT-bbox crop (oracle)
100%
removed
0.944
0.873
Total gain
+9.0 pp
+24.5 pp
from outside-bbox
(mask − full)
+3.2 pp
+11.9 pp
Appendix
Table 11: O2I vs outside-bbox context ablation on INSTRE and ILIAS core (4,715-positive gallery, no 100M distractors; DINOv3-L 768 px, GP queries). BBox attention masking removes outside-bbox information without changing the O2I ratio. The residual gap between the masked and cropped rows isolates the O2I contribution.
Q-enc
G-enc
FwdPass
R@1
R@5
MRR
mAP@1000
856K proxy
full 100M
Vanilla DINOv3-L
vanilla
vanilla
1
54.6
62.1
58.4
25.8
26.5
+ GP (5 scales)
vanilla
vanilla
5
56.3
65.1
60.4
27.2
28.3
+ MSA (3 scales + mask)
vanilla
vanilla
3
57.9
65.9
61.7
27.9
29.2
+ LoRA (rank 8, q,v)
LoRA
LoRA
1
58.1
66.2
61.9
27.9
29.7
Appendix
Table 12: O2I-aware interventions on the 856,089-image ILIAS top-1,000 proxy gallery (1,232 queries, 4,715 positives, 851,374 distractors). Q-enc / G-enc indicate the encoder used for the query and gallery, respectively; FwdPass is the number of query-side forward passes per image. GP and MSA are query-side test-time interventions against the vanilla gallery; the LoRA row re-encodes both query and gallery with the LoRA-adapted weights. Replacing the LoRA gallery with the vanilla one (cross setting) drops the LoRA row to mAP@1000 = 24.4, confirming that the adapter shifts the embedding space and the gallery must be re-encoded to capture the gain. The full 100M column reports the same backbone runs evaluated on the complete ILIAS 100M gallery ( Table 1 ); the proxy underestimates absolute mAP by 0.7–1.3 pp but preserves relative ordering.
Method
mAP@1000
Cost (query)
Cost (gallery)
Pre-reranking (first-stage retrieval)
DINOv3-L (baseline)
26.5
1×
1×
DINOv3-L + GP (5 scales)
28.3
5×
1×
p -value vs. baseline
< 0.001
DINOv3-L + MSA
29.2
3×
1×
p -value vs. baseline
< 0.0001
Appendix
Table 13: Full ILIAS 100M results. DINOv3-L backbone, 1,232 queries, ∼ 100M gallery. p -values are from a paired t -test on per-query mAP@1000 against the DINOv3-L baseline. “Cost (query)” is the online per-query cost; “Cost (gallery)” is the one-time per-image cost amortized over all queries; both in units of one DINOv3-L forward pass at 768 px (1,657 GFLOPs). MSA uses 3 scales with attention masking; GP uses 5 scales without masking. OWLv2 crop reranking uses ∼ 31 detections per image on average (one DINOv3-L crop encoding each, applied once per gallery image and cached); the query-side cost remains a single DINOv3-L forward.
Baseline
MSA Δ
O2I ratio
n
R@1
mAP
Δ R@1
Δ R@5
Δ mAP
<1× (query smaller)
85
74.1
34.6
− 2.4
− 2.4
− 0.7
1 – 3×
479
62.8
33.7
+4.4
+5.4
+2.5
3 – 10×
543
52.3
22.5
+5.5
+7.4
+3.6
>10×
125
23.2
10.9
+2.4
+5.6
+1.8
All
1232
55.0
26.5
+4.2
+5.8
+2.7
Appendix
Table 14: O2I-conditioned gains on ILIAS 100M. MSA improvement over baseline (absolute pp), stratified by query-to-gallery object O2I ratio. Gains are concentrated where the query object is larger than the gallery’s; the small reverse-direction bin ( <1× , query smaller) is slightly degraded.
Paradigm
Model
HF / timm name
Params
Pooling
Self-supervised
DINOv2-L
facebook/dinov2-large
303M
CLS
DINOv2-G
facebook/dinov2-giant
1.1B
CLS
DINOv3-L
facebook/dinov3-vitl16-pretrain-lvd1689m
303M
CLS
Vision-language
CLIP-L (@336)
openai/clip-vit-large-patch14-336
304M
pooled
SigLIP-2-L (@384)
google/siglip2-large-patch16-384
316M
pooled
Supervised
ConvNeXt V2-H
facebook/convnextv2-huge-22k-512
657M
pooled
Appendix
Table 15: Backbone implementation details. Pooling modes: CLS = final-layer class token; pooled = model’s pooled vision-tower output; timm CLS = pre-logit CLS embedding via timm; mean = mean over all patch tokens from the encoder.
Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image regions. However, we show that indiscriminately applying retrieval ignores a critical vulnerability: the resolution-context trade-off. Patch-based zooming recovers details for small targets, but can split large objects and destroy global spatial context; attention-based retrieval better preserves large objects, but remains less reliable on tiny details; and global perception is often fastest when retrieval is unnecessary. Motivated by these failure modes, we introduce ViRGo (Visual Retrieval or Global Perception), a lightweight framework that formulates visual retrieval as an adaptive routing problem. ViRGo estimates object scale from the VLM's intrinsic localization heads during the initial forward pass and combines it with semantic token confidence to select between global perception, patch-based retrieval, and attention-based retrieval with minimal additional computation. Experiments across multiple VQA benchmarks and object-size groups show that ViRGo improves the accuracy-efficiency trade-off: it matches patch retrieval on small details, leverages attention-based retrieval for larger objects, and reduces inference time by routing to the global baseline when zooming is unnecessary.
Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew +2
VinUni-Illinois Smart Health Center, VinUniversity · Texas A&M University
Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve. When the scene contains entities at vastly different scales, existing language-guided generators condition on a single, globally pooled text embedding and quietly drop scale-specific concepts, breaking concept-query retrieval even when pixel fidelity is high. We formalise this failure as semantic collapse and propose CERES, a closed-loop multimodal indexing framework that builds a three-level semantic pyramid, mines implicit concepts via a co-occurrence-aware router, performs scale-routed cross-attention into a lightweight U-Net generator, and verifies coverage by re-indexing the generated image with the same frozen VLM. A continuously differentiable soft-Jaccard coverage objective returns dense gradients to the 0.39 M-parameter generator under explicit non-degeneracy conditions, and coverage is verified by an independent DINOv2 linear probe trained only on external scene and object labels. On four pansharpening benchmarks across seven settings, CERES delivers the new state of the art with the largest gains where scale variation is most extreme (+4.64% relative Q2n and +9.7 mAP for DOTA detection). It also improves concept-query retrieval Recall@5 by +14.0 points and image-text mean reciprocal rank by 0.19 over the strongest baseline, showing that the closed loop preserves queryable content rather than self-referential feature consistency.
Guangyuan Dong, Chuang Liu, Haoyu Wang +10
1National University of Singapore · 2Wuhan University · 4Nankai University +6
No-reference image quality assessment (NR IQA) has recently benefited from deep and multimodal models, yet many SOTA systems still violate at least one basic requirement: they either discard critical quality cues via aggressive resizing, fail to generalize across resolutions, cannot be jointly trained on heterogeneous IQA datasets with mismatched MOS scales, or require prohibitive computation. We present \textbf{ReLIQS}, a model for \textbf{Re}solution-agnostic \textbf{L}earning for \textbf{I}mage \textbf{Q}uality with \textbf{S}aliency, which is resolution-agnostic, preserves original-resolution quality cues, learns from multiple subjective studies, and remains computationally efficient and budget-adaptive. ReLIQS is a CLIP-based multiscale patch-driven architecture that learns both \emph{where to look} and \emph{how to judge} quality. Fixed-size patches are sampled across multiple resolutions, including the original resolution, and encoded with a CLIP vision backbone. A lightweight Perceptual Importance Estimator then predicts IQA-specific importance maps to select a small set of informative patches, and a Latent Quality Axis Module aggregates their embeddings into a single image-level score. Across authentic, synthetic, and AIGC benchmarks spanning diverse resolutions and distortions, ReLIQS generalizes better than strong CNN-, CLIP-, and MLLM-based baselines with matching or reduced computational cost.
Hakan Emre Gedik, Shashank Gupta, Alan Bovik
The University of Texas at Austin · University of Colorado Boulder