Audio Description (AD) makes movies accessible to blind and visually impaired audiences by narrating visual information in gaps between dialogue. Existing automatic AD systems largely treat generation as a local video-to-text problem, assuming that the content to describe and its temporal location are already provided. Realistic AD instead requires coupled decisions about what visual information is narratively important, when it can be spoken without interfering with dialogue, and how it should be formulated to fit within the available time. We formalize AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. A mixed-integer linear program then jointly selects and schedules descriptions across a scene subject to temporal constraints. When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics. Ablations show that explicit temporal constraints drive gains in placement, while salience estimation controls how much narratively useful content is retained. Improvements are concentrated on temporal and narrative measures rather than n-gram overlap, although a significant gap to professional describers remains.
Figures & tables
Figure 1: Excerpt of Audio Description ( italics ) and dialogue ( bold ) from The Girl with the Dragon Tattoo (2011); timestamps mark the start of narration.
Figure 2: Example of the optimization, on the final scene of Figure 1 . Four eventualities occur on screen (top), each with an occurrence midpoint mi , while the soundtrack leaves one gap G between lines of dialogue (middle). The solution (bottom) makes all three decisions jointly: it sets x2=0 , dropping the least salient eventuality because the others cannot otherwise fit; it chooses delivery times di that tile G in order and without overlap; and it selects the compression c4,1 for e4 : the full description c4,0 would exceed the time left, and because the objective weights salience by narration time, the longest variant that fits scores highest. The shaded band shows the Δmax window that keeps a description near the event it describes. Durations and gap boundaries are illustrative.
Gaps
SODA
QEval
C
M
M
T
Acc
T
Expert
51.4
20.1
16.3
81.0
69.6
61.2
Random
1.3
4.4
4.0
27.1
34.1
2.3
DistinctAD
15.6
7.1
8.3
49.0
35.7
8.8
ShotbyShot
16.3
7.0
8.0
53.0
39.9
17.0
Qwen 3.5
13.2
8.1
7.4
38.4
43.9
17.3
Table 1: System performance on REFRAMED challenge set with six evaluation metrics. Indented rows apply the realistic filter. (Dialogue) gap columns are CIDEr and METEOR; SODA columns are SODA-M and SODA-T; QEval columns are QEval (Acc) and QEval-T (T).
SODA-T
QEval
Solve time
LLM
MILP
Δ
LLM
MILP
Δ
< 0.1s
57.5
62.4
+4.9
48.3
51.4
+3.1
0.1–1s
46.9
57.0
+10.1
42.2
45.8
+3.6
1–10s
45.0
55.6
+10.6
43.7
48.3
+4.6
10–60s
37.8
50.1
+12.3
33.9
40.8
+6.9
Table 2: Performance stratified by MILP solve time, an empirical proxy for AD difficulty . As scenes become harder to schedule (longer runtimes), scores drop, while the MILP’s advantage ( Δ ) over the LLM scheduler widens.
Figure 3: References and generated ADs for dialogue gap in Harry Potter and the Goblet of Fire (2005). Dialogue in bold , AD in italics . Multiple-choice question is from the QA-based evaluation.
Gaps
SODA
QEval
C
M
M
T
Acc
T
(1) Qwen MILP
12.9
7.7
7.3
54.9
48.4
28.2
MILP decisions and constraints
(2) w/o grounding
11.0
7.0
7.4
47.8
44.7
16.0
(3) w/o compression
9.2
6.2
7.3
43.7
43.4
22.7
(4) + w/o both
6.6
6.0
7.6
41.0
39.7
12.9
Table 3: Qwen performance on REFRAMED post-production screenplay set using C IDEr and M ETEOR, SODA- M , SODA- T , QEval ( Acc ), and QEval- T .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Occ.
Sal.
Description element
di
e1
0s–4s
0.8
Lisbeth stands with her hands on a table in an upmarket Stockholm tailor shop.
0.0s
e2
0s–4s
0.2
The shop has large windows
–
e3
0s–4s
0.1
that look out onto a road with parked cars,
–
e4
0s–4s
0.1
and there are two mannequin torsos in front of the window.
–
e5
0s–4s
0.3
A tailor is standing behind a large table.
–
e6
0s–4s
0.4
He is a bald white man,
–
Appendix
Table 4: A prose scene description for a scene in The Girl with the Dragon Tattoo (2011). Occurrence is when each description is depicted on screen and salience is a score representing the relevance of each description to the story being told. Occ. is occurrence, and Sal. is salience. Scores are illustrative. di is the optimal delivery start time for selected eventualities; unselected eventualities are marked –.
System
Confusion ( ↓ )
Missed Rate ( ↓ )
Purity ( ↑ )
Coverage ( ↑ )
Baselines
Single speaker
68.1
0.0
31.9
100.0
Changing speaker
95.6
0.0
100.0
4.4
Oracle K
89.0
0.0
33.8
11.7
Open-weight
Pyannote Speaker Diarization 3.1
31.7
0.6
78.0
71.0
Appendix
Table 5: Character speaker diarization performance. Confusion is the proportion of dialogue duration assigned to an incorrect speaker; Missed Rate is the proportion of undetected dialogue. Both are computed after optimal one-to-one mapping between gold and predicted speaker tags. Purity is the proportion of each predicted cluster that belongs to a single gold speaker; Coverage is the proportion of each gold speaker captured by a single predicted cluster. Baselines: Single speaker assigns all dialogue to one tag; Changing speaker assigns a new speaker to every segment; Random Oracle K assigns each segment a random tag drawn from a fixed pool of the true number of speakers. The Final row removes all segments tagged as the AD narrator, where the narrator tag(s) are manually identified and the corresponding segments then automatically discarded (note the slightly higher Missed Rate).
P
R
F1
Baselines
Most frequent character
32.0
32.0
32.0
Random IMDb character
2.9
2.9
2.9
Random reference character
4.9
4.9
4.9
Random-sampled reference character
18.1
18.1
18.1
SDH extraction
Appendix
Table 6: Character speaker identification performance. The task is to identify the IMDb name of the speaker of each dialogue segment. Most frequent character assigns the modal reference character name to every segment; random IMDb character samples uniformly from the full IMDb cast; random reference character samples uniformly from character names in the reference; random-sampled reference character samples them in proportion to their reference frequency. Extracted segments only restricts predictions to segments the system named, leaving the rest unlabelled; speaker-based propagation assigns each speaker in Speechmatics diarization the system-predicted name that wins a majority vote across its segments, then labels every gold segment with the name attached to the diarised speaker it most overlaps with.
Figure 4: Screenplay–movie matching scores (WIP) against a measure of screenplay–movie scene alignment (SIP, defined below) for the 10-movie challenge set. The top row shows SIP metrics derived using human-annotated scene alignments, while the bottom row uses our automatic scene alignment method. The left and middle columns decompose the metrics into their directional components (representing deletion and insertion, respectively).
Figure 11
Figure 7: Distribution of WIP, and its two component parts (left two plots, representing screenplay deletion and insertion respectively), across dataset splits. T=Train, V=Validation, and C=Challenge splits; numbers in brackets represent the number of movies with aligned screenplays in each split.