cs.AISep 29, 2026

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

Authors: Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao

Organizations: New York University · New Jersey Institute of Technology

Abstract

Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

    Jun 19, 2026Sergio Lanza, Jae Hee Lee, Stefan WermterExtraction

  2. Mixture of Cognitive Experts in Large Vision-Language Models

    Jul 12, 2026Robert Wijaya, Ngai-Man CheungRecent Vision-Language ModelsMultimodal Reasoning