What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation
Organizations: AIDAS Laboratory, 1ECE & 2IPAI, Seoul National University
Abstract
Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.
Figures & tables
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
| Role | Checkpoint identifier | Backend | Context limit | Max. output tokens |
| Text | Qwen/Qwen3-30B-A3B-Instruct-2507 | vLLM | 8192 | 1024 |
| ByteDance-Seed/Seed-OSS-36B-Instruct | vLLM | 8192 | 1024 | |
| nvidia/Llama-3_3-Nemotron-Super-49B-v1_5 | vLLM | 8192 | 1024 | |
| Image | Qwen/Qwen3-VL-32B-Instruct | vLLM | 20000 | 1024 |
| OpenGVLab/InternVL3_5-38B-Instruct | vLLM | 20000 | 1024 | |
| LGAI-EXAONE/EXAONE-4.5-33B | vLLM | 20000 | 1024 |
| Statistic | Value | |
|---|---|---|
| Pearson correlation | 0.538 | 0.001 |
| Spearman correlation | 0.704 | 6.9e-06 |
| BERTScore range (0–10) | 8.02–9.08 | – |
| Text LLM-judge range | 3.30–9.39 | – |
| Score paired with DreamSim distance | |||
|---|---|---|---|
| CLIP similarity | -0.982 | 4.0e-23 | 32 |
| Image LLM judge (three-judge mean) | -0.910 | 5.3e-13 | 32 |
| Image modality score | -0.959 | 6.7e-18 | 32 |
| Integrated LLM judge (three-judge mean) | -0.728 | 2.3e-06 | 32 |
| Total Average | -0.809 | 2.1e-08 | 32 |
| Metric paired with integrated judge | |||
|---|---|---|---|
| Text modality score | 0.987 | 3.7e-25 | 32 |
| Image modality score | 0.871 | 8.8e-11 | 32 |
| Speech modality score | 0.870 | 1.0e-10 | 32 |
| DreamSim distance | -0.728 | 2.3e-06 | 32 |
| Partial correlation with integrated judge | Partial | ||
|---|---|---|---|
| Text modality score Image , Speech | 0.921 | 5.7e-13 | 32 |
| Image modality score Text , Speech | 0.637 | 1.5e-04 | 32 |
| Speech modality score Text , Image | 0.123 | 0.519 | 32 |
| Integrated-judge category | Mean | IQR | Loss share | Lowest count |
|---|---|---|---|---|
| Metadata alignment | 8.26 | 1.22 | 18.1% | 0 |
| Multimodal continuity | 7.59 | 1.55 | 25.1% | 0 |
| Condition satisfaction | 7.61 | 2.12 | 24.8% | 0 |
| GT semantic consistency | 6.91 | 1.92 | 32.1% | 32 |
| Integrated-judge category | Text | Image | Speech |
|---|---|---|---|
| Metadata alignment | 0.82 | 0.47 | 0.41 |
| Multimodal continuity | 0.91 | 0.69 | 0.20 |
| Condition satisfaction | 0.93 | 0.53 | -0.12 |
| GT semantic consistency | 0.95 | 0.77 | -0.10 |
| Paradigm | Valid-only | Zero-filled | Change | |
|---|---|---|---|---|
| Orchestration (Large) | 12 | 7.614 | 7.608 | -0.006 |
| Orchestration (Small) | 6 | 6.737 | 6.726 | -0.011 |
| Semi-orchestration (TTS experts) | 4 | 5.428 | 4.719 | -0.709 |
| Semi-orchestration (image experts) | 6 | 6.797 | 6.674 | -0.122 |
| Any-to-any | 4 | 5.792 | 5.613 | -0.179 |
| Experts | Missing candidates | Total Average | ||||||
| Backbone | Image | TTS | T | I | S | Valid-only | Zero-filled | |
| Orchestration (large VLMs) | ||||||||
| GPT-5.4 | D | P | 0 | 0 | 0 | 7.5305 | 7.5302 | -0.0003 |
| GPT-5.4 | K | P | 0 | 0 | 0 | 7.6986 | 7.6983 | -0.0003 |
| GPT-5.4 | D | V | 0 | 0 | 0 | 7.5426 | 7.5397 | -0.0029 |
| GPT-5.4 | K | V | 0 | 0 | 0 | 7.7123 | 7.7094 | -0.0029 |
| Diagnostic | Text | Image | Speech |
|---|---|---|---|
| Weakest-modality count | 9 | 14 | 9 |
| Mean weakest-modality rank gap (pp) | 2.47 | 6.55 | 2.62 |
| IQR of modality score | 1.22 | 1.86 | 0.55 |
| Max Total (Bottom Tercile) | 6.94 | 6.55 | 6.72 |
| Valid-only | Zero-filled | |||
|---|---|---|---|---|
| Association | ||||
| All configurations ( ) | ||||
| Text–Image Speech | 0.626 | 0.560 | 0.001 | |
| Text–Speech Image | 0.758 | 0.812 | ||
| Image–Speech Text | -0.122 | 0.515 | -0.058 | 0.756 |
| Semi-orchestration and any-to-any ( ) | ||||
| Valid-only | Zero-filled | |||
|---|---|---|---|---|
| Modality and controls | ||||
| Text Image, Speech | 0.921 | 0.831 | ||
| Image Text, Speech | 0.637 | 0.378 | 0.039 | |
| Speech Text, Image | 0.123 | 0.519 | 0.342 | 0.065 |
| Track | Submitted | Skipped/unscored | Scored | Category scores |
|---|---|---|---|---|
| Text | 189 | 2 | 187 | 748 |
| Image | 168 | 0 | 168 | 672 |
| Speech | 151 | 0 | 151 | 604 |
| Integrated | 189 | 8 | 181 | 724 |
| Total | 697 | 10 | 687 | 2,748 |
| Track | Spearman | Kendall | Pearson | LOTO range |
|---|---|---|---|---|
| Text | 0.863 | 0.677 | 0.919 | 0.820–0.876 |
| Image | 0.882 | 0.714 | 0.928 | 0.866–0.918 |
| Speech | 0.724 | 0.559 | 0.793 | 0.665–0.800 |
| Integrated | 0.884 | 0.712 | 0.906 | 0.849–0.902 |
| Four-track composite | 0.945 | 0.806 | 0.957 | 0.934–0.955 |
| Paradigm tier | Systems | Human | Three-judge |
|---|---|---|---|
| Orchestration, large VLM | 12 | 7.84 | 7.83 |
| Orchestration, small VLM | 6 | 6.73 | 6.49 |
| Semi-orchestration | 10 | 5.14 | 5.76 |
| Native any-to-any | 4 | 4.61 | 4.90 |
| Track | Human ordinal range | Judge subset/full Spearman |
|---|---|---|
| Text | 0.46–0.72 | 0.946 |
| Image | 0.72–0.84 | 0.966 |
| Speech | 0.26–0.35 | 0.840 |
| Integrated | 0.48–0.64 | 0.932 |
| Score | A–B | A–C | B–C |
|---|---|---|---|
| Text | 0.990 / 0.930 | 0.969 / 0.874 | 0.966 / 0.851 |
| Image | 0.969 / 0.899 | 0.977 / 0.899 | 0.975 / 0.903 |
| Speech | 0.730 / 0.542 | 0.404 / 0.360 | 0.595 / 0.429 |
| Integrated | 0.909 / 0.754 | 0.956 / 0.839 | 0.935 / 0.810 |
| Four-track composite | 0.972 / 0.883 | 0.985 / 0.927 | 0.990 / 0.940 |
| Total Average | 0.986 / 0.919 | 0.993 / 0.948 | 0.992 / 0.948 |
| Track | Pair | Paired observations | Spearman | MAE |
|---|---|---|---|---|
| Text | A–B | 28,194 | 0.873 | 1.011 |
| Text | A–C | 28,191 | 0.872 | 1.017 |
| Text | B–C | 28,191 | 0.888 | 1.433 |
| Image | A–B | 28,177 | 0.864 | 0.994 |
| Image | A–C | 28,177 | 0.806 | 1.078 |
| Image | B–C | 28,177 | 0.803 | 0.933 |
| ID | Backbone | I/S | Total | 95% score interval | Rank | Rank interval |
|---|---|---|---|---|---|---|
| C10 | Claude-Opus-4.7 | K/P | 7.749 | [7.711, 7.786] | 1 | [1, 2] |
| C12 | Claude-Opus-4.7 | K/V | 7.729 | [7.693, 7.765] | 2 | [1, 4] |
| C04 | GPT-5.4 | K/V | 7.712 | [7.676, 7.748] | 3 | [2, 4] |
| C02 | GPT-5.4 | K/P | 7.699 | [7.661, 7.735] | 4 | [2, 4] |
| C08 | GLM-4.6V | K/V | 7.655 | [7.617, 7.691] | 5 | [5, 7] |
| C06 | GLM-4.6V | K/P | 7.631 | [7.591, 7.670] | 6 | [5, 8] |
| Comparison set | Pairs | 95% CI excludes zero | Holm rejections |
|---|---|---|---|
| All pairs | 496 | 476 | 462 |
| Adjacent point-estimate ranks | 31 | 17 | 8 |
| Large orchestration vs. remaining | 240 | 240 | 240 |
| Association | 95% transition interval | ||
|---|---|---|---|
| 32 | Text–Image | 0.824 | [0.815, 0.832] |
| 32 | Text–Speech | 0.880 | [0.855, 0.898] |
| 32 | Image–Speech | 0.693 | [0.661, 0.717] |
| 32 | Text–Image Speech | 0.626 | [0.583, 0.669] |
| 32 | Text–Speech Image | 0.758 | [0.717, 0.787] |
| 32 | Image–Speech Text | -0.122 | [-0.182, -0.054] |
| Component | Range | SD | Share (%) |
|---|---|---|---|
| BERTScore | 8.020–9.081 | 0.153 | 1.62 |
| Text judge | 3.295–9.389 | 1.820 | 28.76 |
| CLIP | 5.827–8.513 | 0.604 | 7.36 |
| Image judge | 2.028–6.808 | 1.623 | 25.29 |
| Speech metadata | 4.190–5.864 | 0.371 | 4.63 |
| Speech judge | 3.286–6.987 | 0.678 | 9.60 |
| Aggregation | First | Top five | ||
|---|---|---|---|---|
| Equal seven | 1.000 | 1.000 | Claude/K/P | 5/5 |
| Modality balanced | 0.995 | 0.972 | Claude/K/P | 5/5 |
| Automatic only | 0.905 | 0.754 | Dynin-Omni | 3/5 |
| Judges only | 0.982 | 0.919 | Claude/K/V | 4/5 |
| Z-score mean | 0.990 | 0.927 | Claude/K/P | 4/5 |
| Min–max mean | 0.993 | 0.940 | Claude/K/P | 4/5 |
| Backbone | I | S | BERT | CLIP | Meta | Total | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.4 | D | P | 0.882 | 0.732 | 0.584 | 9.386 | 5.788 | 6.752 | 8.797 | 7.530 |
| GPT-5.4 | K | P | 0.882 | 0.772 | 0.584 | 9.384 | 6.489 | 6.752 | 8.879 | 7.699 |
| GPT-5.4 | D | V | 0.882 | 0.732 | 0.566 | 9.387 | 5.787 | 6.987 | 8.830 | 7.543 |
| GPT-5.4 | K | V | 0.882 | 0.772 | 0.566 | 9.389 | 6.499 | 6.987 | 8.909 | 7.712 |
| GLM-4.6V | D | P | 0.882 | 0.707 | 0.565 | 9.217 | 5.995 | 6.522 | 8.773 | 7.436 |
| GLM-4.6V | K | P | 0.882 | 0.771 | 0.565 | 9.217 | 6.652 | 6.522 | 8.847 | 7.631 |
| Backbone | Valid | Missing | Mean characters |
|---|---|---|---|
| Ground truth | 900 | 0 | 106.65 |
| GPT-5.4 | 900 | 0 | 226.43 |
| GLM-4.6V | 900 | 0 | 219.19 |
| Claude-Opus-4.7 | 900 | 0 | 244.69 |
| InternVL3.5-4B | 900 | 0 | 240.08 |
| Qwen3.5-4B | 900 | 0 | 146.14 |
| Paradigm | Gender | Speed | Pitch | Emotion | |
|---|---|---|---|---|---|
| Orchestration (Large) | 12 | 85.7 | 58.7 | 48.2 | 33.2 |
| Orchestration (Small) | 6 | 79.8 | 55.4 | 46.5 | 31.1 |
| Semi-orch. (TTS experts) | 4 | 65.4 | 55.7 | 47.0 | 35.9 |
| Semi-orch. (image experts) | 6 | 66.0 | 46.8 | 52.1 | 50.9 |
| Any-to-any | 4 | 66.5 | 49.9 | 48.7 | 38.4 |
| All configurations | 32 | 76.0 | 54.4 | 48.5 | 37.1 |
| Backbone | Image expert | Valid | Missing | Mean | SD |
|---|---|---|---|---|---|
| Dynin-Omni | – | 900 | 0 | 0.2945 | 0.1279 |
| Claude-Opus-4.7 | FLUX.1-Kontext | 900 | 0 | 0.3895 | 0.1046 |
| GLM-4.6V | FLUX.1-Kontext | 900 | 0 | 0.4007 | 0.1114 |
| EMOVA-7B | FLUX.1-Kontext | 876 | 24 | 0.4008 | 0.1212 |
| GPT-5.4 | FLUX.1-Kontext | 900 | 0 | 0.4014 | 0.1108 |
| Qwen3.5-4B | FLUX.1-Kontext | 900 | 0 | 0.4105 | 0.1101 |