Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning
Organizations: Tencent · Peking University · Zhejiang University
Abstract
Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal fusion, but clean evaluation cannot tell whether success depends on stable cross-modal structure or on cues sufficient only in intact inputs. To address this gap, we define a modality fault line: a boundary at which model behavior becomes unstable when a modality remains present and human-interpretable, but its internal evidence structure is perturbed. We introduce SCEval (Structure-Corruption Evaluation) a diagnostic evaluation protocol that keeps the question, answer space, and modality channels fixed while applying controlled structural corruptions to text, vision, and audio individually and jointly. Built from human-verified tri-modal examples from Social-IQ, OmniBench, and VALOR, SCEval evaluates proprietary and open-source omni-modal systems. The results show that structural corruption lowers clean accuracy, text--vision damage forms the most stable shared fault line, and multi-modal degradation is non-additive rather than a simple function of the number of corrupted modalities. Clean omni-modal accuracy therefore does not establish that a model will remain reliable when cross-modal evidence becomes structurally unreliable.
Figures & tables
| Source | Modalities | # Examples |
| SRC Social-IQ | video, audio, text | 100 |
| OmniBench | image, audio, text | 77 |
| VALOR | video, audio, text | 96 |
| Total | — | VERIFY 273 |
| Text | Vision | Audio | |||||||||||||
| Model | Clean | Typo | Drop | Shuffle | Break | Noise | Occ. | LowRes | MBlur | DBlur | Expo. | Bright. | Remove | Mute | Distort |
| Proprietary / API omni-modal models | |||||||||||||||
| Gemini 3.1 Pro | 84.15 | 79.85 4.30 | 72.57 11.58 | 75.54 8.61 | 80.72 3.43 | 73.57 10.58 | 82.88 1.27 | 85.07 –0.92 | 82.29 1.86 | 83.70 0.45 | 85.08 –0.93 | 84.88 –0.73 | 78.89 5.26 | 80.62 3.53 | 86.17 –2.02 |
| Gemini 3 Pro | 82.40 | 76.72 5.68 | 71.30 11.10 | 73.16 9.24 | 80.27 2.13 | 73.00 9.40 | 80.09 2.31 | 83.80 –1.40 | 81.41 0.99 | 82.89 –0.49 | 80.49 1.91 | 83.88 –1.48 | 76.36 6.04 | 75.82 6.58 | 82.51 –0.11 |
| Gemini 3 Flash | 80.95 | 71.06 9.89 | 67.40 13.55 | 70.33 10.62 | 72.89 8.06 | 67.79 13.16 | 80.33 0.62 | 80.97 0.02 | 77.59 3.36 | 78.95 2.00 | 79.37 1.58 | 80.30 0.65 | 74.30 6.65 | 73.24 7.71 | 79.17 1.78 |
| Gemini 3.5 Flash | 78.39 | 72.53 5.86 | 67.03 11.36 | 68.50 9.89 | 73.63 4.76 | 62.75 15.64 | 75.14 3.25 | 78.96 –0.57 | 75.56 2.83 | 77.47 0.92 | 79.30 –0.91 | 78.31 0.08 | 73.45 4.94 | 71.00 7.39 | 77.83 0.56 |
| Operator group | Mean drop (pp) |
| Primary (10 operators) | 5.31 |
| All operators (14) | 5.14 |
| High-rejection stress set (4) | 4.72 |
Appendix figures & tables32 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Modalities | Struct. corruption | Sev. sweep | Joint corruption | Audio/vision variant verif. | Sample-paired protocol | Models eval. | |
| Clean-input omni-modal QA benchmarks | ||||||||
| Social-IQ ( Zadeh et al., 2019 ) | video, audio, text | 7.5k | ✗ | ✗ | ✗ | ✗ | ✗ | — |
| OmniBench ( Li et al., 2026b ) | image, audio, text | 1.1k | ✗ | ✗ | ✗ | ✗ | ✗ | — |
| VALOR ( Liu et al., 2024a ) | video, audio, text | 32k | ✗ | ✗ | ✗ | ✗ | ✗ | — |
| Hallucination / shortcut / distractor suites (clean modalities) | ||||||||
| MMBench ( Liu et al., 2024b ) | image, text | 3.2k | ✗ | ✗ | ✗ | ✗ | ✗ | — |
| Operators | Severities | Gemini 3.1 Pro | Gemini 3 Flash | Qwen3.5-Omni-Plus | MiniCPM-o 4.5 |
| clean | clean | clean | clean | ||
| Text+Vision | canonical = drop_words noise (4 cells) + 2 replacement cells | |||||
| drop_words noise | 76.14 8.01 | 71.55 9.40 | 67.96 5.67 | 63.84 10.66 | |
| drop_words noise | 76.60 7.55 | 70.14 10.82 | 63.01 10.61 | 63.45 11.05 | |
| drop_words noise | 76.47 7.68 | 72.50 8.45 | 70.37 3.26 | 66.47 8.03 | |
| drop_words noise | 77.75 6.40 | 71.43 9.52 | 65.88 7.75 | 65.05 9.45 | |
| Mod. | Operator | Qwen3.5-Omni-Plus | Gemini 3 Flash | ||||||
| s10 | s30 | s50 | s70 | s10 | s30 | s50 | s70 | ||
| Text severity grid | |||||||||
| Text | typo_ocr | 72.91 | 72.93 | 71.16 | 69.60 | 79.40 | 75.40 | 74.04 | 71.06 |
| Text | drop_words | 71.57 | 72.21 | 67.20 | 64.10 | 81.85 | 71.98 | 71.60 | 67.40 |
| Text | word_shuffle | 71.58 | 69.04 | 65.63 | 65.20 | 80.64 | 77.37 | 74.47 | 70.33 |
| Text | sentence_break | 74.33 | 72.54 | 71.27 | 69.60 | 80.82 | 78.03 | 73.33 | 72.89 |
| Model | Group | Valid/expected | Trials | Invalid | Acc. |
| Coverage and invalid-output pressure | |||||
| Qwen3.5-Omni-Plus | Bimodal | 4,368/4,368 | 4,368 | 53 | 67.42 |
| Qwen3.5-Omni-Plus | Trimodal | 6,552/6,552 | 6,552 | 124 | 64.78 |
| Gemini 3 Flash | Bimodal | 4,368/4,368 | 4,368 | 41 | 74.16 |
| Gemini 3 Flash | Trimodal | 6,552/6,552 | 6,552 | 98 | 72.18 |
| Model | Clean | TV | TA | VA | TVA | Panel fault | Strongest cell | |
| Proprietary / API omni-modal models | ||||||||
| Gemini 3.1 Pro | 84.15 | 78.62 | 79.97 | 81.17 | 78.04 | 4.70 | drop_words @sev30 | |
| Gemini 3 Pro | 82.40 | 75.72 | 78.13 | 79.65 | 75.19 | 5.23 | drop_words @sev30 | |
| Gemini 3 Flash | 80.95 | 73.50 | 76.53 | 77.30 | 72.27 | 6.05 | noise @sev70 | |
| Gemini 3.5 Flash | 78.39 | 71.04 | 73.75 | 75.59 | 70.14 | 5.76 | drop_words @sev30 | |
| Gemini 2.5 Pro | 79.12 | 71.92 | 74.71 | 76.50 | 71.03 | 5.58 | drop_words @sev30 | |
| Model | No-text | No-vision | No-audio | Q+Opt | Text-only | VA-only | |||
| Proprietary / API omni-modal models | |||||||||
| Gemini 3.1 Pro | 63.15 | 70.98 | 78.49 | 2.04 | 1.18 | 1.90 | 34.48 | 40.77 | 21.37 |
| Gemini 3 Pro | 67.43 | 67.24 | 78.29 | 2.56 | 0.67 | 2.70 | 32.80 | 43.21 | 29.26 |
| Gemini 3 Flash | 66.24 | 72.01 | 74.59 | 2.84 | 2.78 | 3.79 | 29.36 | 37.70 | 27.38 |
| Gemini 3.5 Flash | 56.46 | 62.76 | 72.17 | 3.21 | 0.68 | 2.39 | 29.34 | 39.06 | 22.79 |
| Gemini 2.5 Pro | 65.68 | 67.89 | 71.52 | 2.96 | 2.25 | 2.97 | 37.01 | 40.54 | 19.06 |
| Model | Std. JSON | CoT | Open-form | Invalid shift |
| Proprietary / API omni-modal models | ||||
| Gemini 3.1 Pro | 85.30 | 83.34 | 80.85 | 2.83 |
| Gemini 3 Pro | 81.59 | 81.68 | 81.79 | 3.81 |
| Gemini 3 Flash | 82.18 | 79.11 | 83.57 | 2.41 |
| Gemini 3.5 Flash | 77.52 | 78.50 | 78.26 | 3.27 |
| Gemini 2.5 Pro | 77.47 | 73.32 | 73.34 | 2.37 |
| Model | Source | Clean | TV | TVA | Strongest drop |
| Proprietary / API omni-modal models | |||||
| Gemini 3.1 Pro | Social-IQ | 85.75 | 81.98 | 77.03 | 8.32 |
| OmniBench | 83.51 | 79.03 | 74.28 | 7.24 | |
| VALOR | 82.76 | 78.72 | 73.99 | 9.73 | |
| Gemini 3 Pro | Social-IQ | 85.62 | 81.44 | 78.33 | 11.37 |
| OmniBench | 78.95 | 72.91 | 69.54 | 7.72 | |
| Comparison | 95% CI | McNemar | Holm | Interpretation | |
| Statistical reliability | |||||
| Clean TV, Qwen3.5-Omni-Plus | 5.97 | [+4.97, +6.97] | <.001 | <.001 | Significant clean-to-fault-line drop |
| Clean TV, Gemini 3 Flash | 7.32 | [+6.32, +8.32] | <.001 | <.001 | Significant clean-to-fault-line drop |
| Weakest single strongest joint, Qwen3.5-Omni-Plus | 0.73 | [-0.27, +1.73] | 0.012 | 0.034 | Tests |
| Weakest single strongest joint, Gemini 3 Flash | 2.45 | [+1.45, +3.45] | <.001 | <.001 | Tests |
| Estimand | Definition | Comparison cohort |
| Confirmatory estimands and cohort alignment | ||
| Paired corruption drop | Same human-valid base examples | |
| Excess model drop | Model drop minus human drop | Same gold-preserved cohort for all terms |
| Asymmetric TV contrast | Common sample seed cohort | |
| Vision / text main effect | Mean at / | Common four-cell cohort |
| TV interaction | Common four-cell cohort | |
| Condition group | Human acc. | Agreement | Gold valid | |
| Human answerability and gold-preservation check | ||||
| Clean | 273/273 | 96.7 | 95.2 | 99.3 |
| Worst text single | 264/273 | 90.5 | 88.6 | 94.7 |
| Worst vision single | 228/273 | 87.3 | 86.4 | 92.4 |
| Worst audio single | 210/273 | 89.6 | 88.2 | 95.1 |
| Strongest TV joint | 198/273 | 80.4 | 79.7 | 88.5 |
| Condition | Sampled gold valid | Model drop on | Human drop | Excess drop | 95% CI |
| Gold-preserved, human-normalized results | |||||
| Worst text single | 94.7% | 10.68 | 1.80 | 8.88 | [7.4, 10.4] |
| Worst vision single | 92.4% | 10.97 | 2.20 | 8.77 | [7.2, 10.2] |
| Worst audio single | 95.1% | 7.02 | 1.50 | 5.52 | [4.1, 6.9] |
| Strongest TV joint | 88.5% | 12.00 | 3.20 | 8.80 | [7.2, 10.4] |
| Strongest TVA joint | 84.8% | 15.50 | 4.60 | 10.90 | [9.0, 12.8] |
| Model | Group | API | Parse | Refusal | Empty | Valid cov. |
| Proprietary / API omni-modal models | ||||||
| Gemini 3.1 Pro | Bimodal | 0 | 3 | 0 | 2 | 266/273 |
| Trimodal | 1 | 4 | 2 | 1 | 263/273 | |
| Gemini 3 Pro | Bimodal | 0 | 4 | 1 | 1 | 266/273 |
| Trimodal | 2 | 1 | 1 | 1 | 265/273 | |
| Gemini 3 Flash | Bimodal | 2 | 1 | 4 | 3 | 259/273 |
| Probe family | Conditions | Metric | Observed result | Claim |
| Mechanism diagnostics | ||||
| Cross-modal mismatch | random, category-matched, answer-matched replacement | 66.74 / 81.19 / 81.96 | Which channel dominates conflict | |
| Temporal order | frame shuffle, frame reversal, middle drop, audio shuffle | temporal drop | 2.76 | Whether media are treated as bags of features |
| AV desync | s, s, s, s, s | desync curve slope | 10.97 | Sensitivity to audiovisual alignment |
| Frame budget | 1, 4, 8, 16, default frames | budget gap | 3.14 | Whether visual fault line is sampling-driven |
| Position bias | key frame first, middle, last | position gap | 9.82 | Primacy/recency in multi-frame prompts |
| Model | Mod. | Weakest cell | Mean acc | Worst-variant | Mean worst | Std (3 var) | |
| Deep-dive models with full single-modality severity grids | |||||||
| Qwen3.5-Omni-Plus | Text | drop_words @sev70 | 64.10 | 62.55 | 1.55 | 0.87 | 3 |
| Vision | occlusion @sev50 | 70.29 | 67.69 | 2.60 | 1.72 | 3 | |
| Audio | mute @sev70 | 69.72 | 67.38 | 2.34 | 1.71 | 3 | |
| Gemini 3 Flash | Text | drop_words @sev70 | 67.40 | 65.69 | 1.71 | 0.88 | 3 |
| Vision | noise @sev70 | 67.79 | 65.23 | 2.56 | 1.31 | 3 | |
| Model | Steepest channel | |||
| Proprietary / API omni-modal models | ||||
| Gemini 3.1 Pro | Text | |||
| Gemini 3 Pro | Text | |||
| Gemini 3 Flash | Text | |||
| Gemini 3.5 Flash | Text | |||
| Gemini 2.5 Pro | Text | |||
| Model family | Mean clean | Steepest channel | ||||||||
| Aggregated by model family (mean across rows, severity 70 single-modality and headline joint cells) | ||||||||||
| Proprietary / API | 6 | 80.07 | 8.33 | 2.34 | 4.58 | 6.99 | 4.65 | 3.06 | 7.94 | Text |
| Open / open-API | 8 | 69.17 | 9.09 | 3.63 | 5.49 | 8.88 | 5.93 | 4.23 | 10.16 | Text |
| Gap (Prop Open) | — | — | ||||||||
| Model | Combo | Mean actual | Mean additive pred. | Mean | % sub-additive | Paired bootstrap | |
| Deep-dive models: canonical joint grids only (replacement cells excluded; see Table 5 ) | |||||||
| Qwen3.5-Omni-Plus | T+V | 4 | 66.81 | 65.87 | 50% | 0.082 | |
| T+A | 4 | 72.26 | 66.47 | 75% | 0.283 | ||
| V+A | 4 | 75.12 | 69.66 | 100% | <.001 | ||
| T+V+A | 8 | 66.76 | 64.18 | 50% | 0.172 | ||
| Gemini 3 Flash | T+V | 4 | 71.41 | 59.90 | 100% | <.001 | |
| Model | ECE (lower is better) | : % wrong answers emitted with high confidence | ||||||
| Clean | Single-TV worst | Joint-TV worst | Joint-TVA worst | Clean | Single-TV worst | Joint-TV worst | Joint-TVA worst | |
| Deep-dive panel calibration | ||||||||
| Qwen3.5-Omni-Plus | 6.40 | 7.49 | 8.98 | 11.12 | 18.6 | 22.0 | 28.1 | 33.7 |
| Gemini 3 Flash | 5.10 | 6.96 | 9.13 | 10.87 | 14.8 | 21.4 | 25.4 | 29.2 |
| Gemini 3.1 Pro | 4.20 | 5.28 | 5.86 | 8.29 | 11.5 | 16.5 | 23.2 | 28.0 |
| MiniCPM-o 4.5 | 7.60 | 9.16 | 10.68 | 11.88 | 23.4 | 25.8 | 31.0 | 35.2 |
| Condition | Mean Jaccard | Shared-hard ( wrong) | Idiosyncratic ( wrong) | Total wrong (union) |
| Failure-set overlap across the 14-model panel | ||||
| Clean baseline | 0.17 | 17/273 | 44/273 | 95/273 |
| drop_words @sev30 (single-T worst) | 0.31 | 59/273 | 47/273 | 172/273 |
| noise @sev70 (single-V worst) | 0.36 | 75/273 | 41/273 | 185/273 |
| mute @sev50 (single-A worst) | 0.28 | 52/273 | 53/273 | 165/273 |
| Joint T+V worst ( drop_words noise @ ) | 0.42 | 103/273 | 35/273 | 210/273 |
| Operator | Annotated | Retained | Retain |
| Audio operators (3) | |||
| mute | 3,324 | 2,362 | 71.1% |
| remove | 3,324 | 2,478 | 74.5% |
| distortion | 1,108 | 1,074 | 96.9% |
| audio subtotal | 7,756 | 5,914 | 76.3% |
| Vision operators (7, image + video unified) | |||
| Mod. | Operator | sev 10 | sev 30 | sev 50 | sev 70 |
| Confound zone: drop rate climbs steeply with severity | |||||
| Vision | occlusion | 0 7.8% | 32.7% | 64.9% | 82.7% |
| Vision | brightness | 0 3.2% | 0 7.2% | 20.9% | 51.6% |
| Audio | mute | 0 6.0% | 18.8% | 35.4% | 55.6% |
| Audio | remove | 0 5.9% | 16.8% | 31.5% | 47.5% |
| Stable zone: drop rate stays low across all severities (confound-free) | |||||
| Rank | Operator | Panel-mean drop (pp) | Group |
| Most damaging severity-70 operators (14-model panel) | |||
| 1 | noise | 11.39 | Primary |
| 2 | drop_words | 11.28 | Primary |
| 3 | word_shuffle | 11.17 | Primary |
| 4 | mute | 7.46 | Coverage-aware stress |
| 5 | remove | 7.20 | Coverage-aware stress |
| Annotation field | Scale | % full agree | Cohen | Fleiss | % adjudicated | ||
| Clean-example annotation (300 candidates) | |||||||
| Clean-example validity | binary | 300 | 2 | 92.4 | 0.83 | — | 7.6 |
| Tri-modal answerability | binary | 300 | 2 | 86.3 | 0.74 | — | 13.7 |
| Modality availability / leakage | 3-way | 300 | 2 | 89.1 | — | 0.71 | 10.9 |
| Failure-reason (rejected items) | 7-way | 27 | 2 | 75.0 | — | 0.62 | 25.0 |
| Corrupted-variant annotation (19,944 audio + vision cells) | |||||||
| Modality | Operator family | Condition key | Structure being damaged | Severity / variants |
| Text structure corruptions | ||||
| Text | Word dropping | drop_words | Removes lexical evidence while preserving the question/options format. | 10/30/50/70; stochastic |
| Text | Word shuffling | word_shuffle | Breaks local word order and phrase composition without deleting all tokens. | 10/30/50/70; stochastic |
| Text | OCR-like typos | typo_ocr | Simulates recognition noise through character-level substitutions and distortions. | 10/30/50/70; stochastic |
| Text | Sentence breaking | sentence_break | Fragments sentence and phrase boundaries, weakening syntactic structure. | 10/30/50/70; stochastic |
| Visual structure corruptions | ||||
| Mod. | Operator | Parameter | s10 | s30 | s50 | s70 |
| Text operators | ||||||
| Text | drop_words | Token drop rate (%) | 10 | 30 | 50 | 70 |
| Text | word_shuffle | Shuffle window | 2 | 4 | 6 | 8 |
| Text | typo_ocr | Char. subst. rate (%) | 10 | 30 | 50 | 70 |
| Text | sentence_break | Fragment prob. (%) | 10 | 30 | 50 | 70 |
| Vision operators | ||||||
| Model | Group | Interface | Clean | Panel fault | Frame-input compatibility |
| Gemini 3.1 Pro | PROP Proprietary | IFACE OK | 84.15 | 4.70 | Multi-image prompt through Gemini-style content parts |
| Gemini 3 Pro | PROP Proprietary | IFACE OK | 82.40 | 5.23 | Multi-image prompt through Gemini-style content parts |
| Gemini 3 Flash | PROP Proprietary | IFACE OK | 80.95 | 6.05 | Multi-image prompt through Gemini-style content parts |
| Gemini 3.5 Flash | PROP Proprietary | IFACE OK | 78.39 | 5.76 | Multi-image prompt through Gemini-style content parts |
| Gemini 2.5 Pro | PROP Proprietary | IFACE OK | 79.12 | 5.58 | Multi-image prompt through Gemini-style content parts |
| Gemini 2.5 Flash | PROP Proprietary | IFACE OK | 75.40 | 6.65 | Multi-image prompt through Gemini-style content parts |
| Panel block | Cells | Condition IDs | Rationale |
| All-model lightweight expansion | |||
| Text+Vision | 4 | bi_tv_drop_words_noise_t30_v30 ; t30_v70 ; t70_v30 ; t70_v70 | Full grid for strongest shared pair |
| Text+Audio | 2 | bi_ta_drop_words_mute_t30_a30 ; bi_ta_drop_words_mute_t70_a50 | Medium/high lexical-acoustic pressure |
| Vision+Audio | 2 | bi_va_noise_mute_v30_a30 ; bi_va_noise_mute_v70_a50 | Non-textual binding check |
| Text+Vision+Audio | 4 | tri_drop_words_noise_mute_t30_v30_a30 ; t30_v70_a30 ; t70_v70_a30 ; tri_word_shuffle_noise_mute_t70_v70_a50 | Representative tri-modal stressors |
| Reported score | 12 | clean + block means + strongest drop + coverage | Comparable fault-line score without running the full 29-cell matrix |
| Combo | Operators | Severities | 29-cell suite | Role | Notes |
| Text+Vision (6 canonical 2 replacement 4 extended = 12 cells) | |||||
| T+V | drop_words noise | yes | canonical | corner of canonical grid | |
| T+V | drop_words noise | yes | canonical | strongest TV cell for both deep-dive models | |
| T+V | drop_words noise | yes | canonical | corner of canonical grid | |
| T+V | drop_words noise | yes | canonical | symmetric heavy-on-both corner | |
| T+V | drop_words occlusion | yes | replacement (vision) | swap vision operator family | |
| Control | Setup | Purpose |
| Task-format and option-shortcut controls | ||
| No-option generation | Remove answer options; the model writes a short answer and evidence. Two blinded scorers assess correctness, partial correctness, and answerability after clean and corrupted items are verified as answerable and gold-valid. | Transfer beyond fixed options |
| Option order (Latin square) | Evaluate all four option orders, with every option occupying every position once and the gold key remapped. Conditions include clean, drop_words @70, noise @70, TV , TV , and TVA . | Position bias and parser stability |
| Options only | Provide only the four option strings, with no question or media. | Answer-text prior |
| Control | Metric | Value |
| Option-shortcut results | ||
| Options only | Panel accuracy | 29.96 |
| Option order (clean) | Mean-order accuracy | 73.16 |
| Option order (clean) | Worst-order accuracy | 70.86 |
| Option order (clean) | Prediction consistency | 94.1% |
| Option order ( noise @70) | Mean–worst-order gap | 3.20 pp |
| Condition | Existing MC drop (pp) | No-option drop (pp) | Transfer (pp) |
| Clean-to-corrupted drop on the shared gold-preserved subset | |||
| drop_words @70 | 11.28 | 13.64 | |
| noise @70 | 11.39 | 14.22 | |
| TV | 10.01 | 12.25 | |
| TV | 8.28 | 9.80 | |
| TVA | 6.36 | 7.95 | |