Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.
Figures & tables
Dataset
Low-level Technique
Mid-level Visual Effect
High-level Narrative Function
CameraBench [ 3 ]
✓
✓
✗
CineTechBench [ 4 ]
✓
✗
✗
ShotBench [ 5 ]
✓
✗
✗
RefineShot [ 6 ]
✓
✗
✗
CinematicVQA
✓
✓
✓
TABLE I: Comparison between CinematicVQA and existing LVLM VQA benchmarks.
Technique Recognition
Visual Presentation
Narrative Function
Multi-Hop Reasoning
Model
Comp
Light
Camera
Comp
Light
Camera
Comp
Light
Camera
Comp
Light
Camera
Avg.
Zero-Shot Prompting
InternVL-3.5-4B
39.16
45.69
38.12
52.74
54.31
59.01
47.52
37.60
37.34
49.87
37.34
47.78
45.54 ±0.12
InternVL-3.5-8B
40.21
45.69
37.34
55.87
54.57
60.05
46.48
41.25
37.34
53.26
45.43
46.21
46.98 ±0.10
Qwen2.5-VL-3B-it
39.86
37.77
34.64
52.39
47.43
52.91
51.61
41.16
33.85
48.47
48.21
39.34
43.97 ±0.16
Qwen2.5-VL-7B-it
43.42
40.29
36.63
56.22
47.86
54.65
46.82
42.38
36.11
45.25
46.82
40.81
44.77 ±0.09
TABLE II: Performance of state-of-the-art LVLMs on CinematicVQA across four reasoning categories: Technique Recognition, Visual Presentation, Narrative Function, and Multi-Hop Reasoning. We highlight the best overall average within each prompting setting in orange, and mark cases where chain-of-thought prompting improves over the corresponding zero-shot result in green.
CoT
Zero-shot
Correct
Incorrect
Total
Correct
1,400
472
1,872
Incorrect
460
2,162
2,622
Total
1,860
2,634
4,494
TABLE III: Paired zero-shot vs. CoT outcomes for Qwen3-VL-8B on the 4,494 questions answerable under both prompts ( 102 are unparseable under CoT and excluded). The off-diagonal flips nearly cancel, and a McNemar test finds no significant difference ( χ12=0.13 , p=0.72 ).
Failure mode
Mechanism
Description displaces evidence
The chain narrates the clip, then matches its own words to the options instead of re-grounding on the footage.
Generic-distractor bias
Broad, neutral wording composes most naturally with the blandest option, which is over-selected.
Negation misalignment
Enumerating what is not seen aligns literally with the foil that recycles a just-written word, over the broader correct option.
Adjective inertia
After an adjective (“stable”, “smooth”) is written, later steps favour the option that reuses it, conflating distinct concepts.
“Trick-question” spiral
With no option matching its description, the model loops in self-doubt and exhausts its budget without committing to a letter.
TABLE IV: Recurring CoT failure modes from the right-to-wrong cases; all share one cause—the model matches its own generated text rather than the video.
Technique Recognition
Visual Presentation
Narrative Function
Multi-Hop Reasoning
Model ( Δ )
Comp
Light
Camera
Comp
Light
Camera
Comp
Light
Camera
Comp
Light
Camera
Avg.
InternVL-3.5-4B (FT)
-1.50
-2.20
-0.80
+4.10
+2.60
+3.30
+4.20
+2.40
+4.80
+4.60
+1.90
+1.80
+2.10
Qwen2.5-VL-3B-it (FT)
-8.68
-2.67
-3.72
+5.94
+0.46
-2.67
+7.25
+2.29
+8.56
+7.78
+6.73
+4.37
+2.14
Qwen2.5-VL-7B-it (FT)
-12.99
-0.99
-6.46
+3.19
-2.29
-3.07
+9.72
+3.71
+7.89
+9.98
+1.62
+7.37
+1.48
Qwen3-VL-4B-it (FT)
-1.22
-2.26
-0.43
+5.31
+3.22
+3.23
+9.75
+5.31
+7.40
+8.70
+6.88
+6.10
+4.33
Qwen3-VL-8B-it (FT)
+1.07
-1.03
-1.81
+7.60
+7.85
+8.63
+10.46
+8.89
+12.03
+13.08
+5.24
+11.77
+6.98
TABLE V: Relative improvement from supervised fine-tuning on CinematicVQA, reported as the change Δ (percentage points) over each model’s Zero-Shot baseline. Green / red denote gains/losses.
Fig. 4: Comparison of base and fine-tuned model performance across different model scales. Bars show performance (accuracy) on the CinematicVQA-eval , with each model’s fine-tuned version plotted adjacent to its base version.
Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is important for fine-grained video understanding and controllable movie-quality video generation, yet remains underexplored in existing multimodal large language models. Unlike question-answering-based evaluation of cinematic understanding, cinematographic captioning requires a unified open-form description over multiple cinematographic dimensions. This task is challenging for two main reasons: the model must infer professional cinematographic concepts from subtle visual evidence, and it must generate captions that are both comprehensive and accurate. Accordingly, we propose CineCap, a framework that combines structured reasoning with spatio-temporal anchors and reinforcement learning with comprehensiveness, accuracy, and gated coverage rewards. The former grounds professional cinematographic descriptions in explicit visual evidence and organizes them into compact atomic reasoning for supervised fine-tuning, while the latter improves the balance between descriptive completeness and factual correctness. In addition, we construct CineCap Bench, a benchmark of 472 manually annotated video-caption pairs for systematic evaluation. Extensive experiments show that CineCap consistently outperforms strong proprietary and open-source baselines, establishing a new state of the art for cinematographic captioning. The code, model checkpoint, and benchmark are publicly available in https://github.com/Hectormxy/CineCap.git.
Xinyu Mao, Yuhui Zeng, Xiaokun Liu +6
The Chinese University of Hong Kong HKSAR · Kling Team, Kuaishou Technology China · National University of Singapore Singapore +1
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Shengyi Wang, Niantong Li, Guangzheng Hu +27
Alibaba Group · Moku Lab, Hujing Digital Media & Entertainment Group · Beijing Film Academy
Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The usefulness of such models have also been demonstrated on a wide range of benchmarks, with an important caveat - the dominant approach in these benchmarks evaluates multiple choice reasoning via text options. This is a natural way to test text-based reasoning in these models, and has led to significant insights regarding model behavior in the community. In this work, we ask a different question - what happens when the evaluation modality is visual, rather than text? We introduce three new vision-centric evaluation benchmarks in temporal frame retrieval, video future prediction, and causal memory distortion, all designed around evaluating visual understanding capabilities in large video models. Our approach complements the existing approaches to evaluate video understanding in frontier models. We show that current frontier models exhibit significant weakness when attempting to reason through visual queries, rather than text. We conclude with an extended analysis section that provides pointers for future improvements in visual understanding for large video models.
Rwiddhi Chakraborty, Yinong, Wang +7
Oliver · University of Copenhagen · Carnegie Mellon University +4