Vision-language models (VLMs) often miss small objects in large images. We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future. We answer them with a theory built on two quantities of the image interface: S, the number of visual tokens across an object's side, and L, the content a call must cover. The limits: a W x H image sent whole within N tokens gives an object of side m at most m*sqrt(N/(WH)) tokens per side. Doubling the token budget N therefore raises the largest S a whole image can reach by only 41%, and seeing an image at the S an object needs costs at least order S^2 tokens whatever the model. If recognition improves gradually with S, any search strategy, zoom agents included, obeys a recall-cost frontier. What better models can change: the S an object needs and how much one call can carry, measured in bits per object found; the coverage cost remains. Decomposition: yes. Assuming only that more tokens per object and less content per call do not hurt on average, splitting an image cannot lower recall if no view zooms out relative to the whole image and views overlap by one object. Neither part of this condition can be dropped; we bound the cost of every such decomposition, and a simple rule approaches the bound as the image grows. We test the theory in about 177,000 requests on 797 images. On controlled images, none of 40 orderings predicted for 8 VLMs is violated. Checked after the fact on drawings, floor plans, natural and synthetic images, 61 of 89 implied orderings hold significantly and 4 fail, all for OpenAI models given more pixels than their default path. On construction drawings the rule never significantly lowered recall relative to the whole image and raised it by up to 0.28. Code and data: https://github.com/shijunzhe/vlm-small-objects
Figures & tables
Figure 1: The two quantities on a construction drawing. (a) A sheet from REDP-X40, our benchmark of 40 construction drawings, with 23 instances of one legend symbol. (b) to (d) One instance as a model with 32 px tokens receives it; blue lines are token boundaries, the orange square is the symbol’s box. Sent whole within a budget of 2,500 tokens, the symbol spans S=0.40 tokens per side (b). A high-resolution mode with four times the budget only doubles S , and the call must still cover and list all 23 symbols (c). One view of the safe rule ( Corollary 1 ) shows the symbol at S=1.5 and contains four symbols (d); its 20 views, each within the same 2,500-token budget, together use 49,000 tokens, 1.3 times the lower bound on the cost of any decomposition that protects the whole sheet at that S ( Theorem 2 (c)).
Experiment
Data
Models
What varies
Runs
Plan
Reported
Configurations, first round
REDP-X40 test (34)
GPT-5.4, Claude Sonnet 4.6, Gemini 2.5, Gemma 4, Qwen3-VL
whole drawing, vendor high resolution, 44 px and model-specific tiles, propose and verify
2
1
§ 5.4 , App. B
Configurations, newer models
REDP-X40 test (34)
GPT-5.5, GPT-5.6, Claude Sonnet 5.5, Claude Opus 5.5, Gemini 3.8, GPT-6.1
whole drawing, tiles
1 or 2
2, 6
§ 5.2 , § 5.4
Synthetic sweeps
156 controlled images
8 model settings
S , number of targets, area, clutter, upsampling
1
3
§ 5.2 , § 5.3
Packed tiles
14 REDP-X40 sheets
GPT-5.6, Claude Sonnet 5.5
tiles per image k=1,2,4,8
1
3
§ 5.3
Reasoning and zoom agents
REDP-X40 test (34)
GPT-5.6, Claude Sonnet 5.5, Gemini 3.8, Qwen3-VL
reasoning; a zoom tool, up to 20 calls
2
3
§ 5.4
Matched reasoning, zoom-tool ablation
REDP-X40 test (34)
GPT-5.6, Claude Sonnet 5.5
one step of the zoom tool at a time; full-resolution call
1 or 2
4
§ 5.4
Table 1: Every experiment at a glance. Runs: repetitions of each model call (the rule reuses the first runs of the whole-image and tile calls). Plan: the registration that fixed the experiment’s data, conditions, and hypotheses before its calls ( Section 4 ). Whole-image and tile conditions of every row also enter the check of Section 5.1 .
(M) crop − whole
(R) 2× crop − crop
(I) native − garbled
(R, M) 2× crop − whole
Thm. 1 : safe tiles − whole
GPT-5.4
+0.12 [ − 0.06, +0.32]
+0.16 [+0.01, +0.29]
+0.43 [+0.28, +0.56]
+0.28 [+0.13, +0.45]
+0.25 [+0.10, +0.40]
GPT-5.6
+0.39 [+0.16, +0.65]
+0.03 [0.00, +0.07]
+0.28 [+0.16, +0.41]
+0.42 [+0.17, +0.69]
+0.37 [+0.15, +0.60]
Claude Sonnet 5.5
+0.17 [ − 0.01, +0.42]
− 0.03 [ − 0.07, +0.01]
+0.49 [+0.26, +0.72]
+0.14 [ − 0.05, +0.38]
+0.09 [ − 0.08, +0.31]
Qwen3-VL
+0.003 [ − 0.15, +0.14]
+0.46 [+0.29, +0.62]
+0.58 [+0.36, +0.79]
+0.46 [+0.25, +0.67]
+0.23 [+0.09, +0.40]
Gemma 4
+0.10 [+0.02, +0.20]
+0.32 [+0.19, +0.47]
+0.19 [+0.10, +0.29]
+0.42 [+0.24, +0.61]
+0.07 [+0.02, +0.13]
Gemini 3.8
+0.29 [+0.11, +0.47]
+0.20 [+0.07, +0.34]
+0.40 [+0.20, +0.62]
+0.49 [+0.34, +0.64]
+0.16 [ − 0.03, +0.35]
Table 2: Direct test of Assumption 1 and Theorem 1 on controlled images. Each cell is the mean difference in hit recall, the fraction of targets detected, over targets interior to both calls, with a 95% bootstrap interval over the twelve parent images; bold: interval above zero; † : added later under an addendum to the registration. Every pair of calls satisfies the definitions of Section 3 exactly, and the tiles satisfy the covering hypothesis of Theorem 1 at Stile=Swhole .
Figure 2: Every single-call ordering implied by Assumption 1 that we found in the data of our other experiments (a post hoc check), grouped by the relation it tests. Each row is a paired difference, predicted better minus predicted worse, with a 95% bootstrap interval, clustered by source drawing where items share one. (a) Theorem 1 as stated: hit recall over targets interior to the whole-image call and to a tile, on items where every tile has at least the whole image’s S ; grey, for comparison, the items where the 44 px recipe lowers S , where the theory makes no prediction and which are not counted. (b) Crops at equal scale (M); more tokens on the same pixels (R); crops at a larger scale, such as a block alone against the same block inside a larger sheet sent at a lower scale (R and M); and the same drawing with more pixels or tokens through a vendor setting (R, with new information where native pixels are sent), on the drawings where the setting adds tokens. The controlled tests of Table 2 are not shown.
Figure 3: Per-drawing whole-drawing F1 on REDP-X40 against the a priori ceiling (34 test drawings; dashed line: observed equals ceiling), for three of the models; Spearman ρ for all of them is given in Section 5.2 . Claude Sonnet 4.6 is read in its documented resized frame.
Figure 4: Controlled stimuli, eight models. (a) Recall against tokens per symbol side S (8 targets and 8 distractors). (b) Recall against the number of targets per call at fixed S (1.25 to 1.43 for resolution-scaled models). (c) Recall against image area with 16 targets of fixed pixel size. Dashed and hollow: first round; solid and filled: second round.
Figure 5: Two generations per vendor under three configurations (strict F1, 34 test drawings; whole drawing read in each model’s documented coordinate convention). Propose and verify: template matching proposes candidates and the model accepts or rejects each one, shown as a crop at the reference’s scale. The grey band spans template matching with the tuning-set and leave-one-out thresholds.
Recall
Difference in recall
Tokens, rule
Cost of rule
Data
Model
Whole
Rule
Rule − whole
Rule − 44 px tiles
/ 44 px tiles
/ Cmin (median)
REDP-X40
GPT-6.1
0.97
1.00
+0.03 [ − 0.004, +0.09]
+0.003 [ − 0.003, +0.01]
1.06
3.5
REDP-X40
Claude Opus 5.5
0.83
0.92
+0.10 [ − 0.004, +0.20]
− 0.05 [ − 0.12, +0.01]
0.61
2.6
REDP-X40
Gemini 3.8
0.62
0.90
+0.28 [+0.15, +0.41]
− 0.04 [ − 0.09, +0.01]
0.20
1.6
REDP-X40
Qwen3-VL
0.37
0.56
+0.19 [+0.06, +0.33]
− 0.07 [ − 0.18, +0.05]
0.69
2.0
FPC-sheets
GPT-6.1
0.63
0.61
− 0.02 [ − 0.07, +0.03]
n/a
n/a
3.3
Table 3: The rule of Corollary 1 at S∗=1.5 against the whole-image call and against the 44 px tile recipe, on the 34 REDP-X40 test drawings and the 60 full FPC-sheets: mean recall per drawing and paired differences with 95% bootstrap intervals over drawings; bold: interval above zero. The last column is the rule’s token cost over the lower bound Cmin of Theorem 2 (c), median over drawings; it is largest where a single view already covers the drawing. n/a: the 44 px recipe was not run on FPC-sheets for GPT-6.1 and Claude Opus 5.5.
Figure 6: (a) Strict F1 on the 14 full sheets with the same medium-effort reasoning throughout. From left to right each bar changes one step: one whole-drawing call; a zoom agent whose tool returns the requested region cut from the overview at the overview’s scale (lower L only); the same crops enlarged to the size of a native crop (also higher S , no new information); native crops (also new information); and, for reference, model-specific tiles. Agent bars average two runs, the overview-crop bars one. Grey band: template matching. (b) The 20 zoom regions GPT-5.6 chose on one sheet (orange), over the overview it received; blue dots are the annotated instances.
Limit
Statement
Support
With current models
Information
Detail lost in the vendor’s resize cannot be recovered; overlays and interpolated zoom add none ( Proposition 1 ). For Claude Sonnet 5.5 native zoom beats enlarged overview crops by +0.09 F1 on full sheets; for GPT-5.6 by only +0.01.
Proposition; ablation
Unchanged; whether it binds depends on the recognizer.
Recognition threshold
Recall rises gradually with tokens per symbol S ; the level at which it rises is a model parameter.
Controlled sweeps; dose response
Differs by model: about one token for Gemini 2.5, 0.3 to 0.5 for Gemma 4, Claude Sonnet 4.6, and Qwen3-VL, at most about 0.3 for GPT-5.4 and the second-round models.
Coverage
Seeing a drawing at S∗ costs at least WHS∗2/(κ2m2) tokens; adaptive zooming pays at least the cost at S∗=Srec with κ=1 , minus area excluded a priori ( Proposition 2 , Proposition 8 ).
Proposition (under Assumption 2 ); agent logs
Formula unchanged; its value scales with Srec2 , which fell, and a zoom agent now pays it on its own.
Per-call capacity
Recall falls with the content of one call: more targets, more area, more tiles per image, or more blocks around a fixed block.
Pre-registered tests; FPC-sheets
Smaller for the newest models, not removed on full drawing sheets.
Interface
Coordinate frame, scale, and format differ by vendor, mode, and even answer; OpenAI high-resolution paths lowered recall although they add pixels; agents add a crop-to-overview conversion, on which two of four agents failed.
Measured
Not improving; must be checked per model and mode.
Verification
With per-item error rate ε , all n items are right with probability at least 1−nε (union bound) and at most 1−ε , and exactly (1−ε)n if errors are independent; so ε≪1/n guarantees exact large counts and, for independent errors, is also needed; locations make results checkable.
Arithmetic
Lowers ε , does not remove the need.
Table 4: Limits on finding many small objects with VLMs, the evidence behind each, and what the newer generation changed.
Recall and F1 fall with tiles per call; k=4 below k=1
recall at k=4 : Claude − 0.13 [ − 0.27, − 0.04], GPT − 0.09 [ − 0.18, +0.01]; F1 at k=4 not significant; not monotone in k ; clear drops at k=8 for both models, and in Claude’s recall already at k=2 and k=4
Not supported as registered
Appendix
Table 7: Third-round hypotheses (registered before the corresponding runs) and outcomes. X7 (annotation audit) was registered without a direction; its result is in Appendix K .
ID
Registered prediction
Result
Verdict
X9-H1
On sheets, the agent beats tiles with the same reasoning (GPT-5.6, Claude Sonnet 5.5)
Supported on FPC-300; on sheets the others also gain slightly
T1-H2
Pooled over budgets, a budget term does not improve a logistic fit of recall on logS
Δ AIC =0.2
Supported
T1-H3
At the same token grid, Gemini 3.8 (low) recalls more than Gemini 2.5
+0.55 [+0.42, +0.67]
Supported
Appendix
Table 8: Fourth-round hypotheses (registered before the corresponding runs) and outcomes. F-H5 and F-H6 were registered on all blocks; first-block values are a post hoc diagnostic (Appendix F ).
Table 9: Input geometry from documentation and from billed tokens.
Resolution
Capacity (fixed S )
Area
Model
S50
R( S<0.5 )
S
R, n=2→96
slope
R, 1 → 8.4 MP ( S )
Gemini 2.5 Flash
1.28
0.11
0.63
0.25 → 0.15
− 0.032
0.41 → 0.01 (0.22)
Gemma 4
0.31
0.35
0.65
1.00 → 0.34
− 0.133
0.88 → 0.02 (0.23)
Qwen3-VL
0.54
0.24
1.25
1.00 → 0.22
− 0.125
0.93 → 0.60 (0.69)
Claude Sonnet 4.6
0.35
0.51
1.43
1.00 → 0.57
− 0.074
0.79 → 0.62 (0.55)
GPT-5.4
< 0.25
0.80
1.25
1.00 → 0.85
− 0.052
0.92 → 0.21 (0.68)
Appendix
Table 10: Controlled stimuli, robust reading. S50 : tokens per symbol side at which recall reaches one half (logistic fit; < 0.25: below 0.25; for GPT-5.4, GPT-5.6, Gemini 3.8, and Claude Sonnet 5.5 it lies at or below the smallest S tested and is extrapolated). R( S<0.5 ): mean recall over images with S<0.5 . Capacity: 1024 2 canvas, symbol 40 px, n targets and n distractors; slope of recall per doubling of n . Area: 16 targets at 40 px on canvases from 1024 2 to 4096 × 2048, with S on the largest canvas in brackets. Upper block: first round; lower block: second round.
Whole image
Tiles (D)
Zoom agent
Model
k=1
2
3
1
2
3
1
2
3
First block only
GPT-5.4
0.61
0.51
0.46
0.36
0.38
0.37
Claude Sonnet 4.6
0.46
0.20
0.11
0.29
0.33
0.34
Gemini 2.5 Flash
0.18
0.11
0.02
0.30
0.21
0.22
Gemma 4 31B
0.38
0.09
0.02
0.57
0.52
0.49
Appendix
Table 11: FPC-sheets, strict F1 for sheets of k×k FloorPlanCAD blocks (1,000 2 to 3,000 2 px), 60 items per area. Empty cells: not run.
Model
Stratum
S (W)
F1 W
F1 U
F1 T
U − W
T − U
neutral prompt (X8b)
GPT-5.6
small
0.40
0.51
0.59
0.73
+0.09 [+0.002, +0.17]
+0.14 [+0.06, +0.22]
medium
1.07
0.83
0.89
0.91
+0.06 [+0.03, +0.10]
+0.02 [+0.01, +0.04]
large
2.41
0.91
0.92
0.93
+0.02 [+0.001, +0.04]
+0.01 [ − 0.01, +0.03]
Claude Sonnet 5.5
small
0.45
0.46
0.53
0.73
+0.08 [+0.03, +0.12]
+0.20 [+0.12, +0.27]
medium
1.22
0.87
0.86
0.90
− 0.02 [ − 0.05, +0.02]
+0.04 [+0.01, +0.07]
Appendix
Table 12: FSC-147, 150 images in three strata of exemplar size ( Section 5.6 ). W: whole image; U: upsampled to a long side of 1,536 px; T: 2 × 2 tiles at the scale of U. S : median tokens per object side in W. Intervals: 95% bootstrap over images.
Model, condition
τsym
ink box
2τsym
GPT-5.4, whole drawing
0.62
0.61
0.67
Claude Sonnet 4.6, whole drawing
0.36
0.33
0.50
GPT-5.4, tiles
0.67
0.67
0.80
Gemma 4, tiles
0.75
0.75
0.76
Qwen3-VL, propose + verify
0.91
0.91
0.91
Claude Sonnet 5.5, tiles
0.92
0.92
0.92
Appendix
Table 13: Strict F1 under alternative matching tolerances.
Model
Condition
miss
dup.
displ.
spur.
GPT-5.4
tiles
0.23
0.02
0.13
0.40
GPT-5.4
propose + verify
0.22
0.02
0.00
0.13
Claude Sonnet 4.6
tiles
0.30
0.04
0.15
0.52
Claude Sonnet 4.6
propose + verify
0.13
0.02
0.00
0.16
Gemini 2.5 Flash
tiles
0.58
0.02
0.25
0.43
Gemini 2.5 Flash
propose + verify
0.12
0.03
0.00
0.22
Appendix
Table 14: Errors per annotated instance (run 0), pooled over all drawings: missed instances, duplicate reports of a found instance, reports within 2τ of a missed instance, and spurious reports elsewhere. Pooled rates need not match the per-drawing F1 averages reported elsewhere.
Model
Condition
calls
input tokens
GPT-5.4
whole drawing
1.0
2.1k
GPT-5.4
tiles
13.4
16.8k
Claude Sonnet 4.6
tiles
13.4
18.5k
Gemma 4
tiles
55.6
52.8k
Gemini 2.5 Flash
tiles
145.8
134.7k
GPT-5.4
propose + verify
57.9
54.1k
Appendix
Table 15: Calls and input tokens per drawing (test set, one run). Template matching runs locally in seconds.
Whole drawing
Vendor
Tiles
Prop. +
Model
A1
A2
hi-res
C
D
verify
GPT-5.4
0.62
0.62
0.51
0.67
0.83
Claude Sonnet 4.6
0.19
0.36
n/a
0.61
0.86
Gemini 2.5 Flash
0.04
0.07
0.18
0.18
0.41
0.85
Gemma 4 31B
0.02
0.31
n/a
0.68
0.75
0.86
Qwen3-VL
0.01
0.01
n/a
0.53
0.91
Appendix
Table 16: First round, strict F1 on the 34 test drawings. A1 reads coordinates as pixels; A2 in each model’s documented convention. For resolution-scaled models C and D are the same configuration. Sheets and views are reported separately in Appendix B .
GPT-5.6
Sonnet 5.5
Gemini 3.8
Whole drawing
0.80
0.84
0.68
Tiles (D)
0.80
0.79
0.91
Propose + verify
0.79
0.80
0.74
Zoom agent
0.94
0.89
n/a
TM, X40 threshold
0.38
TM, oracle threshold
0.54
Appendix
Table 17: REDP-10: 23 drawing and class items with 225 instances from ten further drawings, references from a shared symbol library. Strict F1; no parameter was tuned on these drawings, except the oracle threshold, which is chosen on REDP-10 as an upper bound for template matching.
Model
Whole
Tiles
Agent
GPT-5.4
0.46
0.36
Claude Sonnet 4.6
0.36
0.25
Gemini 2.5 Flash
0.14
0.24
Gemma 4 31B
0.22
0.43
Qwen3-VL
0.29
0.26
GPT-5.6
0.60
0.53
0.74
Appendix
Table 18: FPC-300 (public FloorPlanCAD blocks, 339 test items, 21 classes). Strict F1 with the boundary band ignored. Configurations and prompts are those of REDP-X40, unchanged; template-matching and counter thresholds are set on the 30 tuning items.
Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image regions. However, we show that indiscriminately applying retrieval ignores a critical vulnerability: the resolution-context trade-off. Patch-based zooming recovers details for small targets, but can split large objects and destroy global spatial context; attention-based retrieval better preserves large objects, but remains less reliable on tiny details; and global perception is often fastest when retrieval is unnecessary. Motivated by these failure modes, we introduce ViRGo (Visual Retrieval or Global Perception), a lightweight framework that formulates visual retrieval as an adaptive routing problem. ViRGo estimates object scale from the VLM's intrinsic localization heads during the initial forward pass and combines it with semantic token confidence to select between global perception, patch-based retrieval, and attention-based retrieval with minimal additional computation. Experiments across multiple VQA benchmarks and object-size groups show that ViRGo improves the accuracy-efficiency trade-off: it matches patch retrieval on small details, leverages attention-based retrieval for larger objects, and reduces inference time by routing to the global baseline when zooming is unnecessary.
Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew +2
VinUni-Illinois Smart Health Center, VinUniversity · Texas A&M University
Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored. A natural extension of ``How many r are there in Strawberry?'' asks: how small a visual pattern can a VLM reliably perceive? As such, we introduce FineSightBench, a new benchmark that systematically probes this limit by separating perception tasks (pixel-level recognition of letters, shapes, objects) from reasoning tasks (spatial reasoning, counting, ordering over small targets) across controlled scales of 4--48px. Through comprehensive experiments and detailed failure mode analysis on state-of-the-art models, we reveal a sharp dissociation: perception saturates around 12px, while reasoning remains limited even at larger scales, with persistent numeracy and sequence errors. These findings expose fundamental deficiencies in VLMs' fine-scale visual reasoning that demand more rigorous evaluation.
Lujun Li, Lama Sleem, Niccolo Gentile +4
University of Luxembourg · Foyer S.A. · Université Paris-Saclay
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.
Xinye Zhao, Yunkai Dang, Yunchen Wu +1
School of Intelligence Science and Technology, Nanjing University