cs.CLSep 24, 2026

What, When, and How: Audio Description as Constrained Global Optimization

Authors: Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller

Organizations: School of Informatics University of Edinburgh United Kingdom

Abstract

Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. REFRAMED: Towards Realistic Audio Description Generation for Movies

    Aug 10, 2026Igor Sterner, Mirella Lapata, Alex Lascarides +1Audio UnderstandingFilms

  2. From Visual Cues to Spoken Narration: Rethinking Audio Description

    Sep 1, 2026Akshita Gupta, Aditya Arora, Federico Tombari +2Video CaptioningAudio Understanding

  3. StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

    Jul 13, 2026Seung Hyun Hahm, Minh T. Dinh, SouYoung JinVideo StorytellingNarratives