Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we reveal an individuation bottleneck stemming from image patchification: because Vision Transformers process patches independently, they struggle to group fragmented geometric features across boundaries into distinct object representations. Second, we identify a collapse in the subsequent counting aggregation process, where representation separation rapidly diminishes as numerosity increases due to attention compression. Identifying and formalizing these twin bottlenecks constitutes our first major contribution. To overcome them, we propose ConvStack, a lightweight architecture that operates directly in the visual token space to explicitly aggregate and inject local spatial structures via zero-initialized residual connections. By explicitly addressing the individuation bottleneck, ConvStack provides unambiguous geometric evidence for downstream aggregation. Remarkably, by fine-tuning exclusively on counting tasks, the model achieves substantial improvements in dense object counting and broader spatial understanding benchmarks, without compromising on general visual capabilities.
Figures & tables
Figure 1: Scene understanding versus counting. Qwen3-VL-8B correctly describes the aerial display but counts seven airplanes instead of eight. ConvStack (ours) recovers the correct count.
Figure 2: Local spatial features improve counting. Left: CNN learns faster than ViT on SynPoly. Right: ConvStack nearly closes the accuracy gap between small and large counts. The 1×1 route control leaves a substantial gap.
Figure 3: Patch geometry changes learning speed. Numbers indicate the first epoch with 100% accuracy across all count classes.
Model
Accuracy
MAE
RMSE
Class sep.
ViT
0.779
0.287
0.408
1.570
CNN
0.998
0.027
0.060
2.995
ViT+ConvStack
0.998
0.009
0.037
5.835
Table 1: Counting on SynPoly. Results after 15 epochs with 500 training samples per class. Metric definitions and the evaluation protocol are given in Section B.5 .
Figure 4: ConvStack improves the separation of count classes. t-SNE visualizations of SynPoly image embeddings, colored by object count. ConvStack forms compact, distinct clusters and reduces the class overlap seen in ViT, especially at larger counts.
Figure 5: Separation of count classes tracks counting accuracy. Qwen3-VL-8B on SynPoly: t-SNE of final layer representations and accuracy for each class (left), and accuracy versus hidden state silhouette (right).
Figure 6: Overview of ConvStack. Teal outlines highlight the ConvStack components. Variable names correspond to Equation 7 .
Method
PixMo-Count
CountBenchQA
CV-Bench Spatial
SAT-real
SpatialEval
SAT-Spatial
Models Specialized for Spatial Tasks
Spatial-MLLM-3B
52.67
65.78
75.22
60.00
50.00
58.76
LLaVA-SP Cropping-7B
42.94
45.62
63.59
52.00
31.33
60.61
Honeybee-7B (C-Abstractor)
37.98
54.79
60.50
51.00
28.00
57.63
Qwen3-VL-8B Scale
Qwen3-VL-8B
65.46
89.82
92.34
59.33
61.33
76.79
Table 2: Evaluation on counting and spatial understanding benchmarks. CV-Bench Spatial averages Relation, Depth, and Distance accuracy. The best result is bolded , and the second best is underlined . Dashes denote results that are unavailable or incompatible with the scorer.
Method
POPE
MME
MMBench
MMMU
RealWorldQA
MathVista
Qwen3-VL-8B
87.83
2219.6
88.67
59.00
69.93
66.60
ConvStack-8B
89.52
2251.0
88.60
59.89
70.33
65.00
Table 3: General visual capability. POPE reports F1. Scores are percentages except MME. The better result in each column is bolded.
Configuration
PixMo-Count
SAT-real
ConvStack-8B
75.19
63.33
ViT-stack only
73.09
60.67
LLM-stack only
74.05
61.33
1×1 convolution kernel
73.09
59.33
LoRA
73.85
62.00
Table 4: Ablation study (accuracy %).
Figure 7: Counting accuracy and representations on SynPoly. (a) Accuracy by number class, with 50 images per class. (b,c) t-SNE of the final layer representation of the last prompt token for counts 1 to 10 .
Figure 8: Spatial understanding analyzed with S-Space on 500 MSCOCO images. (a) ConvStack corrects the base model’s depth judgment for a dining table relative to a refrigerator. (b) The corresponding S-Space depth readout shifts from −0.28 (farther) to +0.08 (closer). (c) Aggregate gains over the base model are positive for horizontal, vertical, and depth evaluations.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Parameters (M)
% of base model
Vision layer adapter
0.596
0.007
Shared DeepStack adapter
2.112
0.024
Total added by ConvStack
2.708
0.031
Existing visual merger
40.119
0.458
Total trainable
42.827
0.488
Appendix
Table 5: Parameter budget of ConvStack-8B. Percentages are relative to Qwen3-VL-8B’s 8.767 B parameters, including vision and language components. The shared DeepStack adapter is counted once.
Figure 9: Ablation of the ViT injection layer. Accuracy on PixMo-Count test and SpatialEval for Qwen3-VL-8B with ConvStack at each vision layer. The green band marks layer 18 , selected using PixMo-Count validation accuracy.
Figure 10: Quantitative count separation on SynPoly. (a) Mean cosine silhouette for each of the ten count classes. (b) Separation between counts n and n+1 , relative to their combined spread. Both models are evaluated on the same images. Larger values indicate better separation.
While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cognitive stages: visual individuation, magnitude awareness, and symbolic mapping. Using synthetic Go boards and linear probes, we demonstrate that visual backbones maintain robust, linearly separable representations of quantity well into the extrapolation regime, ruling out perceptual failure. Furthermore, models retain latent magnitude awareness, successfully performing comparative reasoning on quantities they fail to enumerate. We pinpoint the collapse to the symbolic mapping stage, where the model fails to project valid visual magnitudes onto symbolic tokens. Our findings support a frac tured magnitude hypothesis: VLMs fail to acquire a universal number space, instead learning disjoint, modality-specific statistical manifolds that prevent cross-modal grounding for unseen quantities. Validated on the state-of-the-art foundation model, our results suggest that bridging this gap requires inductive priors enforcing unified representations, as data scaling alone is insufficient.
Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by persistent numerical hallucinations. Existing counting benchmarks primarily focus on basic perception in simplified contexts, failing to capture the complex failure modes that emerge under logical constraints or adversarial conditions. To address these limitations, we introduce HoloCount, a holistic and diagnostically rich benchmark structured around a three-level hierarchical taxonomy. HoloCount evaluates MLLMs across: (1) Semantic Counting, focusing on atomic and property-based enumeration; (2) Analytical Counting, assessing logical composition through spatial and set-based reasoning; and (3) Robustness Testing, probing model integrity against adverse scenarios and grounded counter-priors, such as high-density scenes and linguistic biases. Through an exhaustive evaluation of over 20 state-of-the-art MLLMs, we reveal a critical performance gap: even top-tier models degrade significantly as tasks transition from perception to complex analytical reasoning and adverse scenarios. Our findings provide a systematic landscape of current MLLM counting capabilities and offer a roadmap for developing more grounded and reliable multimodal systems. The dataset is available at https://mm-mvr.github.io/HoloCount/.
Multimodal Large Language Models (MLLMs) face a significant inference bottleneck due to the quadratic computational cost of self-attention over long visual token sequences. However, we identify a critical inefficiency in current architectures: Visual Attention Saturation. Our analysis reveals that visual tokens rapidly establish their spatial structure and intra-modal relationships in early layers, rendering visual-to-visual self-attention in deeper layers computationally redundant. Conversely, Feed-Forward Networks (FFNs) in these layers remain essential for projecting visual features into the evolving textual semantic space. Leveraging this insight, we present Visual-Skip (V-Skip), a training-free inference paradigm that decouples spatial interaction from semantic evolution. Rather than discarding tokens, V-Skip imposes block-wise structured sparsity by selectively bypassing saturated visual self-attention modules. Furthermore, recognizing that varying downstream tasks demand distinct reasoning depths, V-Skip employs a lightweight, few-shot calibration to dynamically route the task-optimal sparsity path. Extensive experiments demonstrate that V-Skip effectively bypasses redundant vision attention to achieve block-wise sparsity, maintaining a 94.16% to 100.31% performance retention across diverse MLLMs. Ultimately, we prove that to reason more effectively, models do not need to discard what they see -- they simply need to "look less" at the right depth.