cs.CLSep 30, 2026

Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

Authors: Xia Hu, Brian Potetz, Chun-Ta Lu, Huanfen Yao, Leonidas Guibas, Zhicheng Wang, Howard Zhou, Pengfei Xing, +1 more

Organizations: Google DeepMind · Google Research · Stanford University

Abstract

End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84% of the gain from supplying all prerequisites.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

    Apr 20, 2026Sungeun An, Swanand Ravindra Kadhe, Shailja Thakur +2Reasoning BenchmarkSkills

  2. Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry

    Jun 11, 2026Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3Large Language Models FailAdversarial Robustness

  3. FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses

    Jun 3, 2026Nicholas SabanLarge Language Models FailClustering