ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
Organizations: National Taiwan University · NVIDIA
Abstract
Video-capable vision-language models score above 80% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276
Figures & tables
| Benchmark | QA pairs | Clips | Per-sec. temporalGT | Per-person spatialGT | Multi-person compositional | Iterated human validation |
| VideoMME | 2,700 | 900 | – | – | – | – |
| MVBench | 4,000 | – | – | – | – | |
| EgoSchema | 5,031 | 5,031 | – | – | – | – |
| STAR | 60,000 | 22,000 | – | – | – | |
| PerceptionTest | 11,619 | 691 | ✓ | – | – | |
| TempCompass | 7,540 | 500 | ✓ | – | – | – |
| Model | Transition Sensitivity | Actor Disambig. | Concurrent Binding | Interaction Reasoning | Gaze Detection | Avg. |
| Review subset (1,194 items) | ||||||
| Pooled human reference | 84.5 | 92.4 | 94.0 | 92.8 | 89.6 | 91.0 |
| GPT-5.2 | 68.8 | 76.0 | 76.4 | 70.0 | 47.2 | 67.6 |
| Gemini 3 Flash | 58.3 | 54.4 | 52.4 | 27.6 | 33.2 | 44.6 |
| Gemma-4-31B | 63.9 | 66.0 | 74.4 | 61.6 | 46.0 | 62.3 |
| LLava-Video-32B | 60.3 | 47.2 | 40.8 | 47.6 | 28.4 | 44.1 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Params | Scope |
| Closed-source | ||
| GPT-5.2 OpenAI (2025) | – | Subset |
| Gemini 3 Flash Google DeepMind (2025a) | – | Subset |
| Large open-source ( 27B) | ||
| Qwen3.5-27B Qwen Team (2026b) | 27B | Full |
| InternVL3.5-38B Wang et al. (2025) | 38B | Full |
| Round | Samples | Acc (%) | Primary failure mode | Fix(es) |
| Pre | 281 | 53 | Multiple simultaneous failures | — |
| 1–2 | 54 | 74 | Distractor contamination; auditory actions; degenerate gaze labels; filler distractors | 1–6 |
| 3 | 50 | 74 | Track drift; confusable pairs; shared co-actions; overlapping boxes | 7–12 |
| 4 | 100 | 79 | Boundary detection; crowd confounders | 13 |
| 5 | 100 | 82 | Residual auditory leakage in D1 | 14 |
| 6 | 100 | 84 | Segment centring for D1 | 15 |
| ID | Diagnostic | N | Videos | Duration | People | Options |
| D1 | Transition Sensitivity | 194 | 53 | 4–10 s † | 1 | 4 |
| D2 | Actor Disambiguation | 2,000 | 63 | 5 s | 2–7 | 4 |
| D3 | Concurrent Binding | 2,000 | 64 | 5 s | 1 | 4 |
| D4 | Interaction Reasoning ‡ | 1,611 | 61 | 5 s | 2–7 | 3–4 |
| D5 | Gaze Detection | 896 | 62 | 5 s | 2 | 4 |
| Total | 6,701 | 64 |
| Model | D1 | D2 | D3 | D4 | D5 | WAvg |
| InternVL3.5-8B | 60.3/62.4 (+2.1) | 58.7/39.5 ( 19.2) | 65.5/64.7 ( 0.8) | 59.7/46.5 ( 13.2) | 33.6/33.2 ( 0.4) | 57.6/48.5 ( 9.1) |
| InternVL3.5-38B | 68.0/60.8 ( 7.2) | 65.5/49.8 ( 15.7) | 68.4/65.8 ( 2.6) | 68.7/65.8 ( 2.9) | 38.6/40.4 (+1.8) | 63.6/57.5 ( 6.1) |
| Qwen3.5-4B | 60.3/59.3 ( 1.0) | 66.0/51.4 ( 14.6) | 69.9/67.3 ( 2.6) | 69.0/66.0 ( 3.0) | 35.3/31.6 ( 3.7) | 63.6/57.2 ( 6.4) |
| Qwen3.5-27B | 64.9/64.4 ( 0.5) | 67.7/57.1 ( 10.6) | 74.8/70.9 ( 3.9) | 77.3/79.6 (+2.3) | 43.3/45.7 (+2.4) | 68.8/65.3 ( 3.5) |
| Gemma-4-E4B | 52.1/50.5 ( 1.6) | 47.9/38.1 ( 9.8) | 64.6/64.8 (+0.2) | 43.2/29.7 ( 13.5) | 28.2/27.0 ( 1.2) | 49.2/42.9 ( 6.3) |
| Actor Disambiguation | Interaction Reasoning | |||||
| Model | Box | Coord | Rel | Box | Coord | Rel |
| Qwen3.5-4B | 65.95 | 51.35 | 59.45 | 69.03 | 66.03 | 65.43 |
| Qwen3.5-27B | 67.70 | 57.05 | 62.60 | 77.28 | 79.58 | 73.00 |
| Gemma-4-E4B | 47.90 | 38.10 | 46.75 | 43.20 | 29.70 | 39.79 |
| InternVL3.5-8B | 58.65 | 39.45 | 52.70 | 59.65 | 46.45 | 55.56 |
| InternVL3.5-38B | 65.45 | 49.75 | 59.80 | 68.65 | 65.80 | 64.49 |
| Family (year) | Grounding evidence disclosed | Video/frame handling | Coord. |
| Qwen3.5 (2026) 4/9/27B | Spatial/RefCOCO evaluations; exact grounding-training mix unspecified | Native video processor | |
| InternVL3.5 (2025) 4/8/38B | RefCOCO evaluations and large multimodal SFT; exact grounding-training mix unspecified | 448-pixel frame tiles in our evaluation | |
| LLaVA-Video (2024) 7/32B | Disclosed 178K video set contains captions and QA; box labels not listed | Spatial average pooling, stride 2 in our evaluation | |
| Gemma 3 (2025) 27B | Grounding-task supervision unspecified | 896-pixel image inputs, 256 tokens per image | |
| Gemma 4 (2026) E4B | Grounding-task supervision unspecified; object detection and pointing described | Video as image frames; configurable visual token budget |
| Model | Actor Acc. (%) | Trap rate (%) | Lift |
| VideoLLaMA3-7B | 51.7 | 53.4 | 1.60 |
| LLaVA-Video-7B | 34.0 | 51.9 | 1.56 |
| Gemma-3-27B | 37.9 | 48.3 | 1.45 |
| Qwen3.5-27B | 67.7 | 45.0 | 1.35 |
| LLaVA-OneVision-7B | 47.5 | 44.8 | 1.34 |
| InternVL3.5-38B | 65.5 | 43.1 | 1.29 |
| Trans. Sens. | Actor Disambig. | Conc. Binding | Inter. Reasoning | Gaze Det. | W. Avg. | ||
| Video | InternVL3.5-38B | 68.0 | 65.5 | 68.4 | 68.7 | 38.6 | 63.6 |
| LLava-Video-32B | 60.3 | 45.0 | 42.0 | 51.0 | 28.7 | 43.8 | |
| Gemma-4-31B | 63.9 | 63.0 | 72.9 | 61.6 | 42.0 | 62.8 | |
| Qwen3.5-9B | 64.9 | 64.6 | 69.1 | 70.3 | 37.6 | 63.7 | |
| Qwen3-VL-8B | 61.9 | 57.4 | 70.9 | 66.2 | 30.8 | 60.1 | |
| Text only | InternVL3.5-38B | 69.6 | 43.3 | 40.9 | 43.1 | 23.9 | 40.7 |
| Model | Question micro | Diagnostic macro |
| Qwen3.5-4B | 63.6 [61.0, 66.1] | 60.1 [57.5, 62.5] |
| Qwen3.5-9B | 63.7 [61.3, 66.1] | 61.3 [58.9, 63.5] |
| Qwen3.5-27B | 68.8 [66.4, 71.2] | 65.6 [63.1, 68.0] |
| Qwen3-VL-4B | 59.3 [56.8, 61.9] | 57.8 [55.4, 60.1] |
| Qwen3-VL-8B | 60.1 [57.6, 62.7] | 57.4 [54.8, 60.0] |
| Qwen3-VL-32B | 62.8 [60.0, 65.7] | 60.7 [57.5, 63.7] |
| First model second model | Question micro | Diagnostic macro |
| Qwen3.5-27B Qwen3.5-4B | 5.18 [3.55, 6.68] | 5.52 [3.32, 7.73] |
| Qwen3.5-27B Qwen3.5-9B | 5.06 [3.58, 6.59] | 4.29 [2.20, 6.50] |
| InternVL3.5-38B InternVL3.5-4B | 8.85 [7.00, 10.61] | 8.45 [6.69, 10.15] |
| InternVL3.5-38B InternVL3.5-8B | 5.97 [4.32, 7.60] | 6.30 [4.44, 8.17] |
| LLaVA-Video-32B LLaVA-Video-7B | 3.22 [0.81, 5.65] | 4.23 [1.79, 6.67] |
| Qwen3.5-27B InternVL3.5-38B | 5.18 [3.61, 6.75] | 3.76 [1.69, 5.96] |
| Videos retained | Mean Spearman [95% interval] | Leader retained |
| 8 | 0.946 [0.878, 0.986] | 79.5% |
| 16 | 0.971 [0.926, 0.994] | 96.6% |
| 32 | 0.987 [0.963, 0.998] | 100.0% |
| 48 | 0.993 [0.979, 1.000] | 100.0% |
| Model | Accuracy [95% CI] | Wrong actor / errors |
| Qwen3.5-9B | 57.4 [53.0, 61.7] | 51/213 (23.9%) |
| InternVL3.5-8B | 57.6 [52.0, 63.2] | 34/212 (16.0%) |
| Gemma-4-E4B | 48.0 [43.5, 52.6] | 46/260 (17.7%) |
| Model | Prompt | Trans. | Actor | Concurrent | Interaction | Gaze | Overall |
| Qwen3.5-4B | Direct | 60.3 | 67.2 | 74.0 | 66.8 | 33.6 | 60.4 |
| Step-by-step | 64.9 | 70.0 | 65.2 | 51.6 | 40.8 | 58.2 | |
| Plan-and-solve | 58.8 | 64.0 | 68.0 | 49.6 | 42.4 | 56.4 | |
| Qwen3.5-27B | Direct | 64.9 | 70.0 | 78.4 | 76.4 | 39.6 | 65.9 |
| Step-by-step | 64.9 | 67.6 | 77.2 | 60.8 | 44.4 | 62.9 | |
| Plan-and-solve | 58.8 | 64.4 | 72.8 | 39.6 | 46.0 | 56.2 |
| Model | Blue | Red | Difference [95% CI] |
| Qwen3.5-27B | 67.70 | 67.25 | 0.45 [ 1.58, 0.65] |
| InternVL3.5-8B | 58.65 | 58.65 | 0.00 [ 1.50, 1.50] |
| Gemma-4-E4B | 47.90 | 50.00 | 2.10 [0.30, 4.04] |
| VideoLLaMA3-7B | 51.65 | 49.70 | 1.95 [ 5.12, 0.99] |
| Question micro | Diagnostic macro | |||
| Model | Full | Subset | Full | Subset |
| Qwen3.5-27B | 68.8 | 65.9 | 65.6 | 65.9 |
| InternVL3.5-38B | 63.6 | 61.3 | 61.8 | 61.6 |
| Gemma-4-31B | 62.8 | 62.3 | 60.7 | 62.4 |
| Qwen3.5-4B | 63.6 | 60.4 | 60.1 | 60.4 |
| InternVL3.5-4B | 54.7 | 51.8 | 53.4 | 52.2 |