cs.CVOct 8, 2026

From What to Which: Decoding Modifier Grounding in Frozen MLLMs

Authors: Barbara Toniella Corradini, Caterina Gallegati, Ludovica Genovese, Vittorio Murino

Organizations: AI for Good · AI for Good (AIGO), Istituto Italiano di Tecnologia, Italy · University of Siena, Italy · University of Genoa, Italy · University of Verona, Italy

Abstract

As Multimodal Large Language Models (MLLMs) can describe increasingly complex visual scenes, token-level grounding becomes crucial. Yet, when an MLLM generates "the yellow banana on the left", established grounding approaches focus on what is in the image ("banana"), overlooking tokens that help describe which instance is meant ("yellow", "left"). In this work, we ask whether frozen MLLM representations contain decodable grounding information about the referred instance across generated tokens, extending to modifiers such as attributes, spatial expressions, and relational/action terms. To address this question, we introduce OTTER, a lightweight supervised probe over frozen MLLM representations that uses Optimal Transport (OT) to align generated tokens with visual regions and produce compact grounding maps. Our results show that (i) instance-discriminative visual information can be decoded from modifier tokens, with the clearest evidence for spatial terms, but (ii) is not confined to them, as contextualized object nouns also carry referential information; (iii) the recovered grounding remains informative under context perturbations, while selected regions remain relevant to generation; and (iv) the learned OT-based grounding extends beyond the controlled setting to free generation and cross-dataset transfer.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

    Jul 27, 2026Tianyi Gao, Han Fang, Tianyi Ding +9Multimodal GroundingMultimodal Large Language Models

  2. Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

    Jun 22, 2026Tom Hodemon, Mohamed Chaouch, Aboubacar Tuo +1Visual Question AnsweringVisual Reasoning