Why VLMs Miss Small Objects, and When Zooming In Is Safe
Organizations: Systems Engineering University of California, Berkeley Berkeley, CA, USA · Department of AI Quotr AI Emeryville, CA, USA
Abstract
Vision-language models (VLMs) often miss small objects in large images. We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future. We answer them with a theory built on two quantities of the image interface: S, the number of visual tokens across an object's side, and L, the content a call must cover. The limits: a W x H image sent whole within N tokens gives an object of side m at most m*sqrt(N/(WH)) tokens per side. Doubling the token budget N therefore raises the largest S a whole image can reach by only 41%, and seeing an image at the S an object needs costs at least order S^2 tokens whatever the model. If recognition improves gradually with S, any search strategy, zoom agents included, obeys a recall-cost frontier. What better models can change: the S an object needs and how much one call can carry, measured in bits per object found; the coverage cost remains. Decomposition: yes. Assuming only that more tokens per object and less content per call do not hurt on average, splitting an image cannot lower recall if no view zooms out relative to the whole image and views overlap by one object. Neither part of this condition can be dropped; we bound the cost of every such decomposition, and a simple rule approaches the bound as the image grows. We test the theory in about 177,000 requests on 797 images. On controlled images, none of 40 orderings predicted for 8 VLMs is violated. Checked after the fact on drawings, floor plans, natural and synthetic images, 61 of 89 implied orderings hold significantly and 4 fail, all for OpenAI models given more pixels than their default path. On construction drawings the rule never significantly lowered recall relative to the whole image and raised it by up to 0.28. Code and data: https://github.com/shijunzhe/vlm-small-objects
Figures & tables
| Experiment | Data | Models | What varies | Runs | Plan | Reported |
|---|---|---|---|---|---|---|
| Configurations, first round | REDP-X40 test (34) | GPT-5.4, Claude Sonnet 4.6, Gemini 2.5, Gemma 4, Qwen3-VL | whole drawing, vendor high resolution, 44 px and model-specific tiles, propose and verify | 2 | 1 | § 5.4 , App. B |
| Configurations, newer models | REDP-X40 test (34) | GPT-5.5, GPT-5.6, Claude Sonnet 5.5, Claude Opus 5.5, Gemini 3.8, GPT-6.1 | whole drawing, tiles | 1 or 2 | 2, 6 | § 5.2 , § 5.4 |
| Synthetic sweeps | 156 controlled images | 8 model settings | , number of targets, area, clutter, upsampling | 1 | 3 | § 5.2 , § 5.3 |
| Packed tiles | 14 REDP-X40 sheets | GPT-5.6, Claude Sonnet 5.5 | tiles per image | 1 | 3 | § 5.3 |
| Reasoning and zoom agents | REDP-X40 test (34) | GPT-5.6, Claude Sonnet 5.5, Gemini 3.8, Qwen3-VL | reasoning; a zoom tool, up to 20 calls | 2 | 3 | § 5.4 |
| Matched reasoning, zoom-tool ablation | REDP-X40 test (34) | GPT-5.6, Claude Sonnet 5.5 | one step of the zoom tool at a time; full-resolution call | 1 or 2 | 4 | § 5.4 |
| (M) crop whole | (R) crop crop | (I) native garbled | (R, M) crop whole | Thm. 1 : safe tiles whole | |
|---|---|---|---|---|---|
| GPT-5.4 | +0.12 [ 0.06, +0.32] | +0.16 [+0.01, +0.29] | +0.43 [+0.28, +0.56] | +0.28 [+0.13, +0.45] | +0.25 [+0.10, +0.40] |
| GPT-5.6 | +0.39 [+0.16, +0.65] | +0.03 [0.00, +0.07] | +0.28 [+0.16, +0.41] | +0.42 [+0.17, +0.69] | +0.37 [+0.15, +0.60] |
| Claude Sonnet 5.5 | +0.17 [ 0.01, +0.42] | 0.03 [ 0.07, +0.01] | +0.49 [+0.26, +0.72] | +0.14 [ 0.05, +0.38] | +0.09 [ 0.08, +0.31] |
| Qwen3-VL | +0.003 [ 0.15, +0.14] | +0.46 [+0.29, +0.62] | +0.58 [+0.36, +0.79] | +0.46 [+0.25, +0.67] | +0.23 [+0.09, +0.40] |
| Gemma 4 | +0.10 [+0.02, +0.20] | +0.32 [+0.19, +0.47] | +0.19 [+0.10, +0.29] | +0.42 [+0.24, +0.61] | +0.07 [+0.02, +0.13] |
| Gemini 3.8 | +0.29 [+0.11, +0.47] | +0.20 [+0.07, +0.34] | +0.40 [+0.20, +0.62] | +0.49 [+0.34, +0.64] | +0.16 [ 0.03, +0.35] |
| Recall | Difference in recall | Tokens, rule | Cost of rule | ||||
|---|---|---|---|---|---|---|---|
| Data | Model | Whole | Rule | Rule whole | Rule 44 px tiles | / 44 px tiles | / (median) |
| REDP-X40 | GPT-6.1 | 0.97 | 1.00 | +0.03 [ 0.004, +0.09] | +0.003 [ 0.003, +0.01] | 1.06 | 3.5 |
| REDP-X40 | Claude Opus 5.5 | 0.83 | 0.92 | +0.10 [ 0.004, +0.20] | 0.05 [ 0.12, +0.01] | 0.61 | 2.6 |
| REDP-X40 | Gemini 3.8 | 0.62 | 0.90 | +0.28 [+0.15, +0.41] | 0.04 [ 0.09, +0.01] | 0.20 | 1.6 |
| REDP-X40 | Qwen3-VL | 0.37 | 0.56 | +0.19 [+0.06, +0.33] | 0.07 [ 0.18, +0.05] | 0.69 | 2.0 |
| FPC-sheets | GPT-6.1 | 0.63 | 0.61 | 0.02 [ 0.07, +0.03] | n/a | n/a | 3.3 |
| Limit | Statement | Support | With current models |
|---|---|---|---|
| Information | Detail lost in the vendor’s resize cannot be recovered; overlays and interpolated zoom add none ( Proposition 1 ). For Claude Sonnet 5.5 native zoom beats enlarged overview crops by +0.09 F1 on full sheets; for GPT-5.6 by only +0.01. | Proposition; ablation | Unchanged; whether it binds depends on the recognizer. |
| Recognition threshold | Recall rises gradually with tokens per symbol ; the level at which it rises is a model parameter. | Controlled sweeps; dose response | Differs by model: about one token for Gemini 2.5, 0.3 to 0.5 for Gemma 4, Claude Sonnet 4.6, and Qwen3-VL, at most about 0.3 for GPT-5.4 and the second-round models. |
| Coverage | Seeing a drawing at costs at least tokens; adaptive zooming pays at least the cost at with , minus area excluded a priori ( Proposition 2 , Proposition 8 ). | Proposition (under Assumption 2 ); agent logs | Formula unchanged; its value scales with , which fell, and a zoom agent now pays it on its own. |
| Per-call capacity | Recall falls with the content of one call: more targets, more area, more tiles per image, or more blocks around a fixed block. | Pre-registered tests; FPC-sheets | Smaller for the newest models, not removed on full drawing sheets. |
| Interface | Coordinate frame, scale, and format differ by vendor, mode, and even answer; OpenAI high-resolution paths lowered recall although they add pixels; agents add a crop-to-overview conversion, on which two of four agents failed. | Measured | Not improving; must be checked per model and mode. |
| Verification | With per-item error rate , all items are right with probability at least (union bound) and at most , and exactly if errors are independent; so guarantees exact large counts and, for independent errors, is also needed; locations make results checkable. | Arithmetic | Lowers , does not remove the need. |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Hypothesis | Result | Verdict | |
|---|---|---|---|
| H1 | D beats A1 for every model | Claude +0.43, Gemini 2.5 +0.37, Gemma +0.73, Qwen +0.52 (Holm ); GPT-5.4 +0.05 [ 0.06, +0.17] | Not supported (4 of 5) |
| H2 | Gemma ranks third or lower under A1 and top two under D | fourth under A1 (third under A2), first under D | Supported |
| H3 | D beats C for fixed-grid models | Gemini 2.5 +0.23 [+0.17, +0.30]; Gemma +0.08 [+0.02, +0.14] | Supported |
| H4 | Vendor high resolution does not exceed D | GPT-5.4 0.17 [ 0.30, 0.03]; Gemini 0.24 [ 0.32, 0.14] | Supported |
| H5 | Reading native coordinates (A2) improves localization | Gemma ; Gemini (n.s.); Qwen | Partly |
| H6 | D minus A1 smaller on views than on sheets | holds for GPT-5.4, Claude, Gemini 2.5; fails for Gemma and Qwen | Partly |
| Whole (A2) | Tiles (D) | |||
| Model | sheets | views | sheets | views |
| GPT-5.4 | 0.38 | 0.79 | 0.66 | 0.68 |
| Claude Sonnet 4.6 | 0.14 | 0.52 | 0.59 | 0.63 |
| Gemini 2.5 Flash | 0.01 | 0.11 | 0.42 | 0.41 |
| Gemma 4 | 0.07 | 0.48 | 0.66 | 0.82 |
| Qwen3-VL | 0.00 | 0.01 | 0.52 | 0.53 |
| Hypothesis | Result | Verdict | |
|---|---|---|---|
| X1-H1 | Recall rises with , is near zero below , saturates by 1.5, and curves coincide in | rises for all models; near zero below 0.5 only for Gemini 2.5; ranges from below 0.25 to 1.28 | Partly |
| X1-H2 | At fixed , recall falls with for some models ( not judged) | slope below zero for all six eligible models ( Table 10 ) | Supported |
| X1d-H1 | Recall falls with clutter for at least half the models | no slope interval excludes zero | Not supported |
| X1e-H1 | Upsampling raises recall for resolution-scaled models | GPT-5.6 +0.06 [ 0.01, +0.19], Claude Sonnet 5.5 0.01 [ 0.03, 0.00], GPT-5.4 0.60 (vendor downscaling) | Not supported |
| X1e-H2 | Upsampling leaves fixed-grid models unchanged | Gemini 3.8 +0.03 [ 0.01, +0.08], Gemini 2.5 +0.03 [ 0.01, +0.07], Gemma 0.03 [ 0.10, +0.03] | Supported |
| X2-H1 | Recall and F1 fall with tiles per call; below | recall at : Claude 0.13 [ 0.27, 0.04], GPT 0.09 [ 0.18, +0.01]; F1 at not significant; not monotone in ; clear drops at for both models, and in Claude’s recall already at and | Not supported as registered |
| ID | Registered prediction | Result | Verdict |
|---|---|---|---|
| X9-H1 | On sheets, the agent beats tiles with the same reasoning (GPT-5.6, Claude Sonnet 5.5) | GPT +0.06 [+0.02, +0.12] (Holm ); Claude +0.05 [ 0.005, +0.12] (Holm ) | GPT supported; Claude not |
| X9-H2 | One native-resolution call with reasoning stays below tiles with reasoning on sheets | GPT-5.6 0.58 [ 0.73, 0.43] (Holm ) | Supported |
| X9-H3 | Decomposition A1+R crop only enlarged overview crop native (no direction) | GPT 0.04, +0.11, +0.01; Claude +0.07, +0.06, +0.09 ( Section 5.4 ) | Reported |
| T1-H1 | At fixed pixels, Gemini 3.8 F1 rises from low to default budget where , not elsewhere | FPC-sheets: +0.17 [+0.12, +0.21] ( ), +0.05 [+0.002, +0.09] (others); FPC-300: +0.14 [+0.10, +0.19] and +0.07 [ 0.01, +0.15] | Supported on FPC-300; on sheets the others also gain slightly |
| T1-H2 | Pooled over budgets, a budget term does not improve a logistic fit of recall on | AIC | Supported |
| T1-H3 | At the same token grid, Gemini 3.8 (low) recalls more than Gemini 2.5 | +0.55 [+0.42, +0.67] | Supported |
| Model | Geometry | Measured (input tokens for a blank image) |
|---|---|---|
| GPT-5.4 | 32 px patches, ; 2,500 (auto), 10,000 (original) | : 3,019; original: 11,068 |
| GPT-5.5, GPT-5.6 | 32 px patches, ; 8,192 (auto), 30,000 (5.6 original) | : 4,934; : 9,849; original: 28,219 |
| Claude Sonnet 4.6 | 28 px; 1,568 tokens, long side 1,568 | and larger: 1,542 |
| Claude Sonnet/Opus 5.5 | 28 px; 4,784 tokens, long side 2,576 | : 3,050; : 4,786 |
| Gemini 2.5 Flash | fixed grid, 258 tokens (default) | 258 at every size |
| Gemini 3.8 Flash | fixed grid, 1,090 (default), 2,200 (ultra-high) | 1,092 to 1,111; 2,189 to 2,220 |
| Resolution | Capacity (fixed ) | Area | ||||
| Model | R( ) | R, | slope | R, 1 8.4 MP ( ) | ||
| Gemini 2.5 Flash | 1.28 | 0.11 | 0.63 | 0.25 0.15 | 0.032 | 0.41 0.01 (0.22) |
| Gemma 4 | 0.31 | 0.35 | 0.65 | 1.00 0.34 | 0.133 | 0.88 0.02 (0.23) |
| Qwen3-VL | 0.54 | 0.24 | 1.25 | 1.00 0.22 | 0.125 | 0.93 0.60 (0.69) |
| Claude Sonnet 4.6 | 0.35 | 0.51 | 1.43 | 1.00 0.57 | 0.074 | 0.79 0.62 (0.55) |
| GPT-5.4 | 0.25 | 0.80 | 1.25 | 1.00 0.85 | 0.052 | 0.92 0.21 (0.68) |
| Whole image | Tiles (D) | Zoom agent | |||||||
| Model | 2 | 3 | 1 | 2 | 3 | 1 | 2 | 3 | |
| First block only | |||||||||
| GPT-5.4 | 0.61 | 0.51 | 0.46 | 0.36 | 0.38 | 0.37 | |||
| Claude Sonnet 4.6 | 0.46 | 0.20 | 0.11 | 0.29 | 0.33 | 0.34 | |||
| Gemini 2.5 Flash | 0.18 | 0.11 | 0.02 | 0.30 | 0.21 | 0.22 | |||
| Gemma 4 31B | 0.38 | 0.09 | 0.02 | 0.57 | 0.52 | 0.49 | |||
| Model | Stratum | (W) | F1 W | F1 U | F1 T | U W | T U |
| neutral prompt (X8b) | |||||||
| GPT-5.6 | small | 0.40 | 0.51 | 0.59 | 0.73 | +0.09 [+0.002, +0.17] | +0.14 [+0.06, +0.22] |
| medium | 1.07 | 0.83 | 0.89 | 0.91 | +0.06 [+0.03, +0.10] | +0.02 [+0.01, +0.04] | |
| large | 2.41 | 0.91 | 0.92 | 0.93 | +0.02 [+0.001, +0.04] | +0.01 [ 0.01, +0.03] | |
| Claude Sonnet 5.5 | small | 0.45 | 0.46 | 0.53 | 0.73 | +0.08 [+0.03, +0.12] | +0.20 [+0.12, +0.27] |
| medium | 1.22 | 0.87 | 0.86 | 0.90 | 0.02 [ 0.05, +0.02] | +0.04 [+0.01, +0.07] | |
| Model, condition | ink box | ||
|---|---|---|---|
| GPT-5.4, whole drawing | 0.62 | 0.61 | 0.67 |
| Claude Sonnet 4.6, whole drawing | 0.36 | 0.33 | 0.50 |
| GPT-5.4, tiles | 0.67 | 0.67 | 0.80 |
| Gemma 4, tiles | 0.75 | 0.75 | 0.76 |
| Qwen3-VL, propose + verify | 0.91 | 0.91 | 0.91 |
| Claude Sonnet 5.5, tiles | 0.92 | 0.92 | 0.92 |
| Model | Condition | miss | dup. | displ. | spur. |
|---|---|---|---|---|---|
| GPT-5.4 | tiles | 0.23 | 0.02 | 0.13 | 0.40 |
| GPT-5.4 | propose + verify | 0.22 | 0.02 | 0.00 | 0.13 |
| Claude Sonnet 4.6 | tiles | 0.30 | 0.04 | 0.15 | 0.52 |
| Claude Sonnet 4.6 | propose + verify | 0.13 | 0.02 | 0.00 | 0.16 |
| Gemini 2.5 Flash | tiles | 0.58 | 0.02 | 0.25 | 0.43 |
| Gemini 2.5 Flash | propose + verify | 0.12 | 0.03 | 0.00 | 0.22 |
| Model | Condition | calls | input tokens |
|---|---|---|---|
| GPT-5.4 | whole drawing | 1.0 | 2.1k |
| GPT-5.4 | tiles | 13.4 | 16.8k |
| Claude Sonnet 4.6 | tiles | 13.4 | 18.5k |
| Gemma 4 | tiles | 55.6 | 52.8k |
| Gemini 2.5 Flash | tiles | 145.8 | 134.7k |
| GPT-5.4 | propose + verify | 57.9 | 54.1k |
| Whole drawing | Vendor | Tiles | Prop. + | |||
|---|---|---|---|---|---|---|
| Model | A1 | A2 | hi-res | C | D | verify |
| GPT-5.4 | 0.62 | 0.62 | 0.51 | 0.67 | 0.83 | |
| Claude Sonnet 4.6 | 0.19 | 0.36 | n/a | 0.61 | 0.86 | |
| Gemini 2.5 Flash | 0.04 | 0.07 | 0.18 | 0.18 | 0.41 | 0.85 |
| Gemma 4 31B | 0.02 | 0.31 | n/a | 0.68 | 0.75 | 0.86 |
| Qwen3-VL | 0.01 | 0.01 | n/a | 0.53 | 0.91 | |
| GPT-5.6 | Sonnet 5.5 | Gemini 3.8 | |
|---|---|---|---|
| Whole drawing | 0.80 | 0.84 | 0.68 |
| Tiles (D) | 0.80 | 0.79 | 0.91 |
| Propose + verify | 0.79 | 0.80 | 0.74 |
| Zoom agent | 0.94 | 0.89 | n/a |
| TM, X40 threshold | 0.38 | ||
| TM, oracle threshold | 0.54 | ||
| Model | Whole | Tiles | Agent |
|---|---|---|---|
| GPT-5.4 | 0.46 | 0.36 | |
| Claude Sonnet 4.6 | 0.36 | 0.25 | |
| Gemini 2.5 Flash | 0.14 | 0.24 | |
| Gemma 4 31B | 0.22 | 0.43 | |
| Qwen3-VL | 0.29 | 0.26 | |
| GPT-5.6 | 0.60 | 0.53 | 0.74 |