CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles
Organizations: University of Alberta · Independent Researcher
Abstract
Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6,072 tracks and 8.13M timestamped subtitle entries. It provides a matched multi-task, multilingual, and multicultural (MultiX) evaluation setting: seven tasks span narrative reconstruction and abstraction, genre prediction, age suitability, country-specific motion-picture ratings across ten national classification systems, and subtitle-grounded language safety. Across nine LLMs, plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies substantially across models and languages; national rating systems expose distinct calibration patterns; and strong profanity is far easier to ground than mild obscenity. CineSubBench establishes film as a long-context LLM evaluation domain and provides a unified benchmark for measuring narrative, multilingual, cultural, and evidence-grounding capabilities.
Figures & tables
| Rank | Model | Family | Composite | Narrative Understanding and Generation | Cultural Prediction and Assessment | LC Assessment | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Overall score (0–100) | Plot | Synopsis | Key message | Overall | Genre micro-F1 (%) | Age exact (%) | Age 1 (%) | Age MAE | Country EM (%) | Country MAE | LC agnostic F1 (%) | LC strict F1 (%) | |||
| Gemini 3.8 Flash | Closed frontier | 78.04 | 4.16 | 3.34 | 3.11 | 3.54 | 84.93 | 34.25 | 81.42 | 0.93 | 77.87 | 0.31 | 86.8 | 77.8 | |
| GPT-5.6 Sol | Closed frontier | 74.7 | 4.1 | 3.4 | 3.08 | 3.52 | 81.55 | 31.08 | 78.75 | 0.98 | 65.42 | 0.44 | 81.7 | 74.5 | |
| DeepSeek V4 Pro 1.6T | Open-weight | 65.52 | 3.91 | 2.93 | 2.9 | 3.25 | 81.63 | 30.67 | 79.33 | 0.98 | 58.03 | 0.52 | 57.5 | 48.3 | |
| 4 | Gemini 3.5 Flash Lite | Closed baseline | 64.7 | 3.83 | 2.91 | 2.99 | 3.25 | 80.92 | 31.83 | 72.17 | 1.09 | 64.65 | 0.46 | 65.3 | 42.8 |
| 5 | Claude Haiku 4.5 | Closed frontier | 57.5 | 3.45 | 2.44 | 2.84 | 2.91 | 71.45 | 27.42 | 72.75 | 1.16 | 54.32 | 0.57 | 58.1 | 37.2 |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Primary role | Quality and provenance treatment |
|---|---|---|
| Kids-in-Mind | Initial 6,186-film inventory; review-derived language metadata. | The native identifier supports cross-source traceability; records without a usable IMDb link for subtitle retrieval are excluded. |
| Common Sense Media | General age-suitability label, storyline, and review page. | Cross-linked through IMDb ID; anomalous embedded IMDb links receive case-by-case manual adjudication before retention. |
| IMDb | Stable title identifier; plot synopsis, runtime, film metadata, and regional motion-picture ratings. | Provides the entity key used to link sources and subtitle assets; populated plot-synopsis and rating fields are mandatory for the metadata-complete cohort. |
| Subtitle providers | Timestamped SRT files in six languages, with provider metadata. | OpenSubtitles is the primary source and SubDL supports recovery. The highest-download candidate is selected per film and language when alternatives exist; its count is retained in the JSON. Structural and temporal verification remains independent of this popularity proxy. |
| Step | Decision | Retention rule | Films |
|---|---|---|---|
| 1 | Source crawl | Kids-in-Mind films collected | 6,186 |
| 2 | IMDb link screen | 4,307 films with links; exclude 33 invalid/search-unusable entries | 4,274 |
| 3 | Entity linking | IMDb-ID agreement across Kids-in-Mind and Common Sense Media; anomalous links manually reviewed | 2,831 |
| 4 | Metadata completeness | Populated IMDb plot synopsis and motion-picture-rating fields | 2,322 |
| 5 | Language coverage | Subtitles in all six selected languages | 1,231 |
| 6 | Country coverage | Ratings from all ten selected countries in the final benchmark | 1,012 |
| Country | Ordered rating labels |
|---|---|
| Australia | G, PG, M, MA15+, R18+ |
| Brazil | Livre, 10, 12, 14, 16, 18 |
| France | Tous publics, 12, 16, 18 |
| Germany | 0, 6, 12, 16, 18 |
| Netherlands | AL, 6, 9, 12, 14, 16, 18 |
| Singapore | G, PG, PG13, NC16, M18, R21 |
| Verification layer | Automated rule | Audit purpose and response |
|---|---|---|
| Provider download count | Among alternative SRT files for the same film and language, select the file with the highest source-reported count; retain the selected count in the JSON. | Use popularity only as an initial quality proxy. Subject the selected file to the remaining checks and recover it if verification finds a genuine problem. |
| Provider-body integrity | Scan subtitle bodies for server-error signatures, including Cloudflare and Guru Meditation messages. | Detect error pages returned in place of dialogue; retrieve or inspect a replacement asset. |
| Timestamp syntax | Require valid SRT time ranges with hours, minutes, seconds, millisecond separator, and arrow structure (e.g., 00:09:09,949 --> 00:09:12,179 ). | Catch malformed headers before JSON creation; repair or re-download the source asset. |
| Chronological consistency | Compare the start and end time of every subtitle block. | Flag negative-duration temporal inversions. |
| Single-block lifespan | Flag any subtitle block lasting 15 minutes or more. | Detect frozen or malformed subtitle segments. |
| Delayed onset | Flag a first operational subtitle beginning after five minutes. | Detect missing opening dialogue or late-start tracks. |
| Tier | Encoding | Purpose and representative case |
|---|---|---|
| 1 | utf-8-sig | Default modern-web decoding while removing a byte-order mark when present. |
| 2 | utf-16 | Handles UTF-16 subtitle assets, including files exported by automated transcription or speech-to-text systems. |
| 3 | cp1256 | Handles legacy Windows Arabic and Persian encodings, preserving right-to-left dialogue instead of dropping unsupported characters. |
| 4 | cp1252 | Handles legacy Western-European encodings and extended diacritics, such as é , ü , and ~n . |
| Failure | none succeeds | Reject the asset and require manual normalization or a clean replacement before ingestion resumes. |
| Source website | Public dataset fields |
|---|---|
| Common Sense Media | ageSuitabilityRating ; storyline . |
| Kids-in-Mind | keyMessage ; languageContentReference . |
| IMDb | title ; releaseYear ; duration ; imdbRating ; countriesOfOrigin ; originalLanguages ; genres ; topics ; plot ; synopsis ; motionPictureRatings ; directors ; writers ; topCast ; fullCast . |
| OpenSubtitles and SubDL | subtitles , including each track’s language , downloadCount , and timestamped entries ( index , timeframe , and content ). |
| Cine Sub Bench annotation pipeline | languageContentAssessment , containing manually audited, subtitle-localized category counts and evidence indices derived from the English subtitle track. |
| Statistic | Value |
|---|---|
| Films | 1,012 |
| Subtitle tracks | 6,072 |
| Subtitle entries | 8.13M |
| Release years | 1977–2026 |
| Runtime, mean / median | 114.6 / 113.0 min |
| Runtime, 95th pct. / max | 149.4 / 242.0 min |
| Language | Mean tok. | Median tok. | 95th pct. | Max tok. | Median entries | Median ratio |
|---|---|---|---|---|---|---|
| English (en) | 43,123 | 41,759 | 68,833 | 144,268 | 1,418 | 1.00 |
| Arabic (ar) | 39,867 | 38,536 | 65,023 | 98,701 | 1,214 | 0.95 |
| Indonesian (id) | 40,641 | 39,739 | 64,760 | 120,034 | 1,308 | 0.97 |
| Persian (fa) | 45,162 | 43,916 | 73,731 | 112,920 | 1,325 | 1.08 |
| Romanian (ro) | 41,010 | 39,676 | 65,542 | 129,799 | 1,244 | 0.99 |
| Vietnamese (vi) | 43,613 | 42,374 | 70,575 | 108,761 | 1,322 | 1.05 |
| Reference field | Mean | Median | P95 | Max |
|---|---|---|---|---|
| Plot | 33 | 33 | 50 | 66 |
| Synopsis | 1,448 | 1,262 | 2,802 | 7,151 |
| Storyline | 156 | 154 | 242 | 437 |
| Key message | 14 | 12 | 26 | 55 |
| Reference field | Mean | Median | P95 | Max |
|---|---|---|---|---|
| Plot | 33 | 33 | 50 | 66 |
| Synopsis | 1,448 | 1,262 | 2,802 | 7,151 |
| Storyline | 156 | 154 | 242 | 437 |
| Key message | 14 | 12 | 26 | 55 |
| Category | Gold count | Evidence lines |
|---|---|---|
| Strong profanity | 21,892 | 20,602 |
| Crude bodily language | 20,131 | 19,149 |
| Religious profanity/exclamation | 10,861 | 10,447 |
| Mild obscenity | 10,360 | 10,088 |
| Language | Flagged | Span/run. | Span p05 | Tok. ratio | Entry ratio | Low tok. | Low entry |
|---|---|---|---|---|---|---|---|
| English (en) | 9 | 0.95 | 0.89 | 1.00 | 1.00 | 0 | 0 |
| Arabic (ar) | 142 | 0.93 | 0.87 | 0.95 | 0.88 | 27 | 120 |
| Indonesian (id) | 112 | 0.95 | 0.89 | 0.97 | 0.93 | 28 | 84 |
| Persian (fa) | 123 | 0.96 | 0.88 | 1.08 | 0.95 | 23 | 77 |
| Romanian (ro) | 155 | 0.95 | 0.89 | 0.98 | 0.89 | 27 | 125 |
| Vietnamese (vi) | 93 | 0.96 | 0.90 | 1.04 | 0.95 | 20 | 57 |
| Genre | Films |
|---|---|
| Drama | 531 |
| Comedy | 380 |
| Thriller | 350 |
| Adventure | 340 |
| Action | 336 |
| Fantasy | 215 |
| Genre | Films |
|---|---|
| Drama | 531 |
| Comedy | 380 |
| Thriller | 350 |
| Adventure | 340 |
| Action | 336 |
| Fantasy | 215 |
| Age label | Films |
|---|---|
| 16+ | 157 |
| 17+ | 153 |
| 14+ | 130 |
| 13+ | 129 |
| 15+ | 109 |
| 12+ | 66 |
| Evaluation role | Models | Model access | Subtitle inputs | Tasks |
|---|---|---|---|---|
| Closed frontier | GPT-5.6 Sol ; Gemini 3.8 Flash ; Claude Haiku 4.5 | OpenAI, Gemini, and Anthropic APIs | English + Arabic, Indonesian, Persian, Romanian, Vietnamese | Narrative and cultural; English language safety where available |
| Closed baseline | GPT-5 Nano ; Gemini 3.5 Flash Lite | OpenAI and Gemini APIs | English + Arabic, Indonesian, Persian, Romanian, Vietnamese | Narrative and cultural; English language safety where available |
| Open-weight | DeepSeek V4 Pro 1.6T ; Qwen3 235B ; Mistral 4 119B ; Llama 4 Scout 17B | OpenRouter API | English + Arabic, Indonesian, Persian, Romanian, Vietnamese | Narrative and cultural; language safety for completed English lanes |
| Local scaling | Ministral 3B , 8B , and 14B | Local GPU inference | English | Controlled English scaling analysis |
| Narrative evaluator | GPT-5.6 Luna | OpenAI API | Narrative outputs from all compared models | Reference-based narrative scoring and structured error annotations |
| Task | Paired scores | Pearson | Spearman | Weighted | Exact (%) | Within 1 (%) | Mean |
|---|---|---|---|---|---|---|---|
| All three tasks | 1350 | 0.808 | 0.815 | 0.778 | 56.8 | 98.4 | |
| Plot | 450 | 0.761 | 0.710 | 0.675 | 43.4 | 96.9 | |
| Synopsis | 450 | 0.750 | 0.751 | 0.729 | 61.5 | 98.2 | |
| Key message | 450 | 0.766 | 0.776 | 0.763 | 65.5 | 100.0 |
| Model | Plot (S/T) | Synopsis (S/T) | Key message (S/T) | Mean change [95% CI] |
|---|---|---|---|---|
| GPT-5.6 Sol | 4.58 / 4.68 | 3.38 / 3.54 | 3.00 / 2.96 | [ , ] |
| Gemini 3.8 Flash | 4.58 / 4.68 | 3.58 / 3.36 | 3.14 / 3.20 | [ , ] |
| Claude Haiku 4.5 | 4.12 / 3.94 | 2.48 / 2.24 | 2.94 / 2.70 | [ , ] |
| DeepSeek V4 Pro | 4.38 / 4.52 | 3.30 / 2.96 | 2.98 / 3.06 | [ , ] |
| Llama 4 Scout | 2.82 / 2.60 | 1.88 / 1.60 | 2.50 / 2.30 | [ , ] |
| Model | Plot (S/C) | Synopsis (S/C) | Key message (S/C) | Mean change [95% CI] |
|---|---|---|---|---|
| GPT-5.6 Sol | 4.70 / 3.80 | 3.30 / 2.70 | 3.00 / 3.05 | [ , ] |
| Gemini 3.8 Flash | 4.65 / 4.50 | 3.55 / 2.80 | 3.15 / 3.00 | [ , ] |
| Claude Haiku 4.5 | 4.10 / 3.95 | 2.40 / 2.35 | 2.90 / 2.75 | [ , ] |
| DeepSeek V4 Pro | 4.45 / 4.25 | 3.20 / 2.60 | 2.95 / 3.15 | [ , ] |
| Llama 4 Scout | 2.80 / 2.65 | 1.90 / 1.70 | 2.55 / 2.25 | [ , ] |
| Model | Micro-F1 (S) | Micro-F1 (T) | Change (pp) | Exact (S, %) | Exact (T, %) |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 75.7 | 82.3 | 20.0 | 30.0 | |
| Gemini 3.8 Flash | 79.8 | 84.7 | 24.0 | 40.0 | |
| Claude Haiku 4.5 | 67.3 | 72.4 | 14.0 | 8.0 | |
| DeepSeek V4 Pro | 79.3 | 81.8 | 30.0 | 34.0 | |
| Llama 4 Scout | 61.2 | 72.2 | 2.0 | 18.0 |
| Model | Composite score | Difference from Gemini (95% CI) | Holm-adjusted |
|---|---|---|---|
| Gemini 3.8 Flash | 78.04 | Reference | — |
| GPT-5.6 Sol | 74.70 | ||
| DeepSeek V4 Pro 1.6T | 65.52 |
| Model | Family | Narrative Understanding and Generation | Cultural Prediction and Assessment | ||||||||||
| Overall (1–5) | Genre micro-F1 (%) | Age exact (%) | Country EM (%) | ||||||||||
| EN | CL | EN | CL | EN | CL | EN | CL | ||||||
| GPT-5.6 Sol | Closed frontier | 3.56 | 3.52 | -0.05 | 82.2 | 81.4 | -0.8 | 33.5 | 30.6 | -2.9 | 67.2 | 65.1 | -2.1 |
| Gemini 3.8 Flash | Closed frontier | 3.61 | 3.52 | -0.09 | 84.7 | 85.0 | +0.3 | 36.5 | 33.8 | -2.7 | 78.8 | 77.7 | -1.1 |
| Claude Haiku 4.5 | Closed frontier | 3.02 | 2.89 | -0.14 | 72.2 | 71.3 | -0.9 | 29.5 | 27.0 | -2.5 | 58.0 | 53.6 | -4.4 |
| GPT-5 Nano | Closed baseline | 2.66 | 2.47 | -0.19 | 74.8 | 72.0 | -2.8 | 19.5 | 20.9 | +1.4 | 43.8 | 43.9 | +0.1 |