cs.CVAug 30, 2026

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

Authors: Kaishuu Shinozaki-ConefreyOlivier PascaudRobin CourantXi WangDimitris SamarasVicky Kalogeiton

Organizations: LIX, Ecole Polytechnique, IP Paris, Palaiseau, France · New York University, New York, NY, USA · Ecole nationale supérieure Louis-Lumière, Saint-Denis, France · Stony Brook University, Stony Brook, NY, USA

Abstract

Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal large language models (MLLMs) are evaluated almost exclusively on understanding what happens rather than why it is presented that way. We introduce TAKE 85, the first benchmark for directorial-intent understanding, comprising 398 short films (85 hours) with expert-verified question-answer pairs spanning global and fine-grained visual and audio intent. Through controlled modality ablations, TAKE 85 enables systematic evaluation of multimodal reasoning. Experiments on state-of-the-art MLLMs reveal a substantial gap between perceptual recognition and intentional understanding: while models accurately describe events and narratives, they consistently fail to infer the communicative role of filmmaking decisions. Our results establish directorial intent as a previously overlooked dimension of multimodal understanding: even the strongest model reaches only 58 out of 100, and our ablations show that no input modality is sufficient on its own. All code, Q&As, and models are publicly available from https://github.com/KaiShinozakiConefrey/Take-85

Explore similar work

CardsList
  1. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    Jul 27, 2026Shengyi Wang, Niantong Li, Guangzheng Hu +27FilmsAi-Generated

  2. Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

    Aug 13, 2026Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera +1Audio-Visual ReasoningSocial Media