NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning
Organizations: New York University · New Jersey Institute of Technology
Abstract
Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.
Figures & tables
| CV-Bench | BLINK | Other Benchmarks | ||||||||
| Model | Overall | Count | Depth | Dist. | Rel. | Overall | MV. | Loc. | RealWorldQA | MMStar |
| Representative VLM backbones | ||||||||||
| DeepSeek-VL1 Lu et al. (2024) | 61.6 | 59.0 | 63.2 | 58.2 | 68.5 | 38.1 | 50.4 | 37.7 | 50.5 | 38.9 |
| Idefics3-8B-Llama3 Laurençon et al. (2024) | 67.7 | 60.5 | 72.8 | 67.2 | 73.4 | 42.7 | 45.9 | 50.8 | 62.0 | 49.3 |
| Phi-4 Multimodal Abouelenin et al. (2025) | 71.7 | 68.5 | 74.2 | 70.8 | 75.1 | 49.7 | 48.1 | 56.6 | 61.8 | 59.7 |
| InternVL3-8B Zhu and others (2025) | 81.3 | 70.9 | 84.8 | 83.1 | 89.7 | 51.3 | 51.1 | 58.2 | 68.2 | 59.4 |
| CV-Bench | BLINK | Other Benchmarks | ||||||||
| Model | Overall | Count | Depth | Dist. | Rel. | Overall | MV. | Loc. | RealWorldQA | MMStar |
| Vision-token reduction methods | ||||||||||
| FastV Chen et al. (2024a) | 75.9 | 64.0 | 83.5 | 74.5 | 85.5 | 48.4 | 55.6 | 49.2 | 68.8 | 55.1 |
| SparseVLM Zhang et al. (2025b) | 69.8 | 54.3 | 74.7 | 68.8 | 86.3 | 46.9 | 55.6 | 59.0 | 50.9 | 53.8 |
| PruMerge Shang et al. (2025) | 71.6 | 55.2 | 78.0 | 74.0 | 83.2 | 46.8 | 54.9 | 51.6 | 64.3 | 50.6 |
| MustDrop Liu et al. (2024b) | 73.7 | 58.6 | 81.5 | 74.8 | 83.9 | 47.2 | 55.6 | 52.5 | 65.1 | 53.0 |
| CV-Bench | BLINK | |||||||
|---|---|---|---|---|---|---|---|---|
| Setting | Overall | Count | Depth | Dist. | Rel. | Overall | MV. | Loc. |
| Baseline (Qwen2.5-VL-7B) | 78.5 | 67.1 | 87.2 | 76.2 | 87.2 | 52.1 | 55.6 | 53.3 |
| Baseline + FT | 76.5 0.3 | 66.6 0.2 | 86.7 0.3 | 69.8 0.9 | 86.5 0.6 | 51.1 0.1 | 55.6 1.9 | 55.7 0.4 |
| + Dense Cross-Attn + PCS | 79.4 0.3 | 66.2 0.6 | 87.0 0.4 | 79.5 1.2 | 88.8 0.3 | 51.3 0.3 | 55.6 1.9 | 54.9 1.8 |
| + Random SNS + NVF | 78.6 0.4 | 67.6 0.5 | 86.5 0.6 | 76.6 1.0 | 86.4 0.6 | 51.6 0.3 | 55.6 1.5 | 54.9 1.0 |
| + PCS only | 78.6 0.4 | 67.3 0.4 | 86.8 0.3 | 76.8 1.4 | 87.2 0.6 | 51.3 0.2 | 55.6 1.9 | 54.9 0.4 |
| CV-Bench | BLINK | |||||||
|---|---|---|---|---|---|---|---|---|
| Layer | Overall | Count | Depth | Dist. | Rel. | Overall | MV. | Loc. |
| 4 | 80.5 0.1 | 67.3 0.2 | 87.0 0.6 | 83.0 0.4 | 88.5 0.3 | 51.6 0.3 | 57.2 1.2 | 56.6 1.1 |
| 8 | 81.6 0.3 | 67.9 0.6 | 87.0 0.4 | 85.7 0.8 | 89.5 0.5 | 52.8 0.2 | 63.9 0.4 | 56.6 0.9 |
| 12 | 78.3 0.4 | 65.2 0.8 | 85.7 0.8 | 79.5 1.2 | 86.3 0.3 | 51.3 0.2 | 55.6 0.1 | 54.1 0.1 |
| 16 | 79.0 0.3 | 67.7 0.6 | 87.2 0.2 | 76.9 0.5 | 87.7 0.6 | 51.8 0.5 | 56.9 1.9 | 56.3 0.5 |
| 20 | 78.6 0.4 | 66.8 0.3 | 86.8 0.4 | 77.0 1.1 | 86.9 0.6 | 51.7 0.1 | 55.6 1.9 | 55.7 1.3 |
| CV-Bench | BLINK | |||||||
|---|---|---|---|---|---|---|---|---|
| Clusters | Overall | Count | Depth | Dist. | Rel. | Overall | MV. | Loc. |
| 32 | 80.4 0.2 | 68.4 0.5 | 86.1 0.4 | 82.1 1.2 | 88.2 0.5 | 52.0 0.2 | 57.2 1.5 | 53.9 2.6 |
| 64 | 81.6 0.3 | 67.9 0.6 | 87.0 0.4 | 85.7 0.8 | 89.5 0.5 | 52.8 0.2 | 63.9 0.4 | 56.6 0.9 |
| 128 | 80.3 0.1 | 66.4 1.1 | 87.1 0.4 | 82.8 0.7 | 88.1 0.4 | 51.7 0.3 | 57.8 1.7 | 58.3 1.3 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Hyperparameter | Value |
|---|---|---|
| Sparse Neuron Space | Expansion ratio | 32 |
| Sparsity level | 32 | |
| Sparsity coefficient | 0.05 | |
| Min. image threshold | 3 | |
| Neuron-guided Visual Focus | Top- patches per cluster | 60 |
| Top- clusters | 5 |
| Evaluation Setting | Base Qwen2.5-VL | NeuronEye | Difference |
|---|---|---|---|
| Overall Accuracy | 65.30% | 63.10% | -2.20 |
| Anatomy Identification | 48.52% | 48.10% | -0.42 |
| Disease Diagnosis | 63.74% | 60.94% | -2.80 |
| Lesion Grading | 57.61% | 55.43% | -2.18 |
| Modality Recognition | 97.57% | 96.18% | -1.39 |
| Other Biological Attributes | 71.72% | 66.21% | -5.51 |
| Configuration | Latency (s) | Lat. (%) | Peak Mem. (MB) | Mem. (%) |
|---|---|---|---|---|
| Base (Qwen2.5-VL) | 0.382 | – | 16,857 | – |
| + SNS + NVF | 0.456 | +19.4 | 23,933 | +42.0 |
| + SNS + NVF + PCS | 0.457 | +19.7 | 23,933 | +42.0 |
| Module | Component | Params | % of Base | Stage 1 | Stage 2 |
|---|---|---|---|---|---|
| SNS | SAE | 822.08M | 10.79% | Train | Frozen |
| Neuron Clusters | – | – | Built | Frozen | |
| NVF | Router | 6.54M | 0.09% | Train | Train |
| Vision-Side Scorer | 6.54M | 0.09% | Train | Train | |
| Projection Layers | 12.85M | 0.17% | – | Train | |
| Localized Cross-Attention | 51.39M | 0.67% | – | Train |