cs.CVSep 28, 2026

ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models

Authors: Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timothée Lardy, Hung-Ting Su, Winston H. Hsu

Organizations: National Taiwan University · NVIDIA

Abstract

Video-capable vision-language models score above 80% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

    May 2, 2026Alejandro Aparcedo, Akash Kumar, Aaryan Garg +5Fine-Grained Video UnderstandingSpatio-Temporal Reasoning

  2. FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding

    May 19, 2026Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh +2Fine-Grained Video UnderstandingRecent Vision-Language Models

  3. SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

    May 29, 2026Tianhui Liu, Jie Feng, Zhiheng Zheng +6Spatial ReasoningRecent Vision-Language Models