cs.AISep 27, 2026

When Does the Concept of "Dog" Emerge in an Audio LLM?

Authors: Zhe Wang, Shiqi Liu, Ruiyun Zhong, Tiechong Zhu, Yihua Tan

Organizations: Huazhong University of Science and Technology, Wuhan, China

Abstract

Multimodal large language models answer audio questions, but how they represent auditory semantics and use them in decisions remains unclear, limiting our understanding of response formation. We study dog barking in Qwen2.5-Omni-7B using Jacobian lens (J-lens) readout and directional interventions. We define the dog direction as a J-lens-derived hidden-state vector associated with dog; adding or removing its component modulates dog-related information. We find this information decodable without dog/bark prompt cues or animal-identification requirements. Directional interventions change response tendencies and some final answers, with effects concentrated in late-layer states immediately before generation across species classification, vocalization classification, and sound description. The dog direction shows no comparable advantage over controls in animal/other classification. These results provide causal-intervention evidence that the dog direction affects output scores in a task-dependent manner, most consistently at L22 and L24 immediately before generation.

Figures & tables

Explore similar work

CardsList
  1. Learning When to Think While Listening in Large Audio-Language Models

    May 26, 2026Zhiyuan Song, Weici Zhao, Yang Xiao +3Large Audio Language Models

  2. Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs

    Sep 29, 2026Yueran Ma, Ronghao LinOmni-ModalCross-Modal