End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84% of the gain from supplying all prerequisites.
Figures & tables
Figure 1 : An example CADET task ( left ) decomposed into 6 unit tasks formalized as an SCM ( middle ). Right : natural capability (NC) of each unit task ( top ), and their causal contribution ( N / S(1,∅) ; bottom ) to final task-6, where supplying unit task-2 alone raises final accuracy from 0.27 to 0.65.
Contrast
N-Score (Success drop from ⋯ )
S-Score (Failure recovery from ⋯ )
(1,0)
corrupting Xi into a wrong answer in oracle env.
fixing Xi into the correct answer in all-wrong env.
(1,∅)
leaving Xi unassisted in oracle env.
providing the correct Xi alone in natural env.
(∅,0)
injecting a wrong answer for Xi in natural env.
removing a wrong answer for Xi in all-wrong env.
(0,∅)
removing Xi ’s wrong answer in all-wrong env.
injecting a wrong Xi as a scaffold in natural env.
Table 1: Operational semantics of N-Score, S-Score across 4 ordered state contrasts (a,a′) .
Figure 2 : Overview of SCMs corresponding to 10 tasks in CADET.
Figure 3 : Natural(NC) and intrinsic(IC) capabilities on non-root unit tasks: (a) category-level NC and IC averaged over 12 models; (b–e) per-model NC and IC within each category.
Figure 6 : N/S(1,∅) across Pattern 1( a,b ) and Pattern 2 ( c,d ) tasks for top-4 SOTA ( a,c ) and remaining 8 Others ( b,d ) models ( sqrt(⋅) -scaled axes, similar in Fig 5 ).
Figure 6Figure 7
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Task Name
Modality
Category
# Instances
Answer Format
Music Note Counting
Image
Perception
114
Integer
Dice Sum
Image
Perception
130
Integer
Map Navigation
Image
Spatial
107
Integer
Maze
Image
Spatial
80
Multiple-Choice
Surround Localization
Multi-Image
Spatial
154
Integer
Multi-View Sequencing
Multi-Image
Temporal
154
Letter List
Appendix
Table 2 : Overview of the 10 overall evaluation tasks in CADET.
Task
#Units
#Roots
Depth
∣Pa(Y)∣
Category
#Questions
Has Markov
Music Note Counting
4
2
2
3
P/C
1,163
No
Dice Sum
4
2
2
3
P/C
1,919
No
Map Navigation
3
1
2
2
P/S/C
918
Yes
Maze
2
1
1
1
P/S
880
Yes
Surround Localization
6
3
2
5
P/S
4,928
No
Multi-View Sequencing
6
2
3
5
P/S/T
4,158
No
Appendix
Table 3 : Structure of the task-level SCM of each evaluation task. – on multistep reasoning is due to it start from a Markov chain structure.
Node
Unit Task
Cat.
Answer Format
#Q
Parents
Music Note Counting
X1
Object Identification
P
Multiple-choice
821
—
X2
Object Classification
P
Multiple-choice
114
X1
X3
Completeness Verification
P
Multiple-choice
114
—
Y
Object Counting
C
Integer
114
X1,X2,X3
Dice Sum
Appendix
Table 4 : Unit task specification of the 10 evaluation tasks in CADET. Y denotes the overall task outcome; ↻ marks a Markov chain unit task.
Figure 11 : Prompt template for our evaluation under active causal interventions ( Ai∈{1,0} ). Under the natural state ( Ai=∅ ), only the target question and last paragraph of output format instructions are prompted.
Figure 12 : LLM-as-a-judge prompt template.
Figure 13 : Comparing NC on root unit tasks against NC and IC on non-root unit tasks, per category, averaged over the 12 models. Cognitive is not shown as it has no root unit task. The non-root bar stacks IC − NC on top of NC, so its top is IC.
Figure 14 : Comparing NC on root unit tasks against NC and IC on non-root unit tasks for each of the 12 models. There are 16 root unit tasks and 30 non-root unit tasks. The non-root bar stacks IC − NC on top of NC, so its top is IC. Models are ordered by overall non-root NC.
Figure 15 : Comparing NC and RC across the four categories on non-root unit tasks, averaged over the 12 models. Each bar shows NC with the signed NC/RC difference stacked out of it, so the boundary of the shaded band is RC.
Figure 16 : Comparing NC and RC for each of the 12 models within each category on non-root unit tasks. Circles mark NC and triangles mark RC; the shaded band spans the two and is colored by the sign of RC − NC. The horizontal lines are the means over the 12 models.
#unit tasks
Root
Non-root
Category
root
non-root
NC
NC
IC
IC − NC
Perception
11
4
0.775 ± 0.077
0.590 ± 0.138
0.796 ± 0.112
+0.206
Spatial
3
9
0.627 ± 0.100
0.556 ± 0.077
0.671 ± 0.059
+0.115
Temporal
2
8
0.714 ± 0.119
0.460 ± 0.072
0.614 ± 0.087
+0.154
Cognitive
0
9
–
0.406 ± 0.092
0.721 ± 0.089
+0.316
Appendix
Table 6 : NC on root unit tasks and NC, IC on non-root unit tasks, per category, with the number of unit tasks in each group. Entries are the unweighted mean over unit tasks, averaged over the 12 models, ± the standard deviation across models.
Model
Root NC
Non-root NC
Non-root IC
NC(root) − NC(non-root)
NC(root) − IC(non-root)
GPT-5.4
0.781 ± 0.108
0.612 ± 0.180
0.791 ± 0.167
+0.169
− 0.010
Gem 3.5F
0.839 ± 0.099
0.595 ± 0.217
0.783 ± 0.172
+0.244
+0.056
Gem 3.1P
0.842 ± 0.090
0.577 ± 0.202
0.781 ± 0.173
+0.265
+0.062
Gem 3F
0.811 ± 0.119
0.528 ± 0.206
0.757 ± 0.180
+0.282
+0.053
Qwen27B
0.780 ± 0.132
0.515 ± 0.247
0.711 ± 0.219
+0.264
+0.069
Qwen397B
0.749 ± 0.144
0.487 ± 0.240
0.647 ± 0.226
+0.262
+0.101
Appendix
Table 7 : NC on root unit tasks and NC, IC on non-root unit tasks for each of the 12 models. Entries are the unweighted mean over unit tasks ± the standard deviation across unit tasks (16 root, 30 non-root). The last two columns are differences of the corresponding means.
Category
NC
IC
RC
IC − NC
RC − NC
Perception
0.590 ± 0.138
0.796 ± 0.112
0.446 ± 0.131
+0.206
− 0.144
Spatial
0.556 ± 0.077
0.671 ± 0.059
0.439 ± 0.102
+0.115
− 0.116
Temporal
0.460 ± 0.072
0.614 ± 0.087
0.414 ± 0.063
+0.154
− 0.046
Cognitive
0.406 ± 0.092
0.721 ± 0.089
0.445 ± 0.060
+0.316
+0.040
Appendix
Table 8 : NC, IC and RC per category on the 30 non-root unit tasks. Entries are the unweighted mean over unit tasks, averaged over the 12 models, ± the standard deviation across models.
Perception
Spatial
Temporal
Cognitive
Overall
Model
NC
IC
RC
NC
IC
RC
NC
IC
RC
NC
IC
RC
NC
IC
RC
GPT-5.4
0.636 ± 0.194
0.913 ± 0.087
0.507 ± 0.288
0.665 ± 0.202
0.747 ± 0.214
0.549 ± 0.213
0.560 ± 0.148
0.730 ± 0.181
0.537 ± 0.219
0.593 ± 0.192
0.835 ± 0.094
0.575 ± 0.274
0.612 ± 0.180
0.791 ± 0.167
0.548 ± 0.232
Gem 3.5F
0.722 ± 0.192
0.889 ± 0.042
0.599 ± 0.194
0.672 ± 0.231
0.754 ± 0.213
0.572 ± 0.252
0.564 ± 0.154
0.699 ± 0.202
0.481 ± 0.187
0.490 ± 0.232
0.839 ± 0.091
0.507 ± 0.219
0.595 ± 0.217
0.783 ± 0.172
0.532 ± 0.212
Gem 3.1P
0.761 ± 0.155
0.907 ± 0.027
0.678 ± 0.158
0.650 ± 0.217
0.740 ± 0.230
0.584 ± 0.207
0.515 ± 0.164
0.698 ± 0.175
0.450 ± 0.222
0.478 ± 0.174
0.839 ± 0.088
0.513 ± 0.194
0.577 ± 0.202
0.781 ± 0.173
0.540 ± 0.206
Gem 3F
0.664 ± 0.250
0.879 ± 0.072
0.439 ± 0.304
0.570 ± 0.268
0.710 ± 0.253
0.431 ± 0.240
0.524 ± 0.122
0.707 ± 0.163
0.422 ± 0.205
0.430 ± 0.154
0.796 ± 0.121
0.450 ± 0.223
0.528 ± 0.206
0.757 ± 0.180
0.435 ± 0.222
Qwen27B
0.682 ± 0.167
0.873 ± 0.119
0.418 ± 0.237
0.571 ± 0.308
0.663 ± 0.252
0.439 ± 0.249
0.515 ± 0.248
0.672 ± 0.210
0.466 ± 0.276
0.386 ± 0.159
0.722 ± 0.220
0.421 ± 0.213
0.515 ± 0.247
0.711 ± 0.219
0.438 ± 0.233
Appendix
Table 9 : NC, IC and RC per model and category on the 30 non-root unit tasks, with Overall aggregating all 30. Entries are the unweighted mean over unit tasks ± the standard deviation across unit tasks.
Figure 17 : Prerequisite N/S -Scores under state contrasts (a) (1,0) , (b) (0,∅) , and (c) (∅,0) , averaged across models. Dashed circles mark the highest- S prerequisite within each task. Axes and quadrant thresholds follow Figure 5 .
Figure 18 : Prerequisite N/S -Scores under (1,0) (a–d), (0,∅) (e–h), and (∅,0) (i–l) by task pattern and model group. Large markers denote group means; grey markers indicate individual models.
Figure 19 : Pairwise Spearman rank correlation between models on N -Scores (a–d) and S -Scores (e–h) across prerequisites under the four state contrasts. Lines separate SOTA from Others.
Figure 20 : Prerequisite N/S -Scores under the four state contrasts, averaged over evaluated models.
Figure 25
Figure 23 : Per-model N - and S -Scores for the parents of Map Nav. under the four state contrasts. The vertical line separates SOTA from Others.
Figure 24 : Same as Figure 23 , for Dice Sum.
Figure 25 : Same as Figure 23 , for Music Note.
Figure 26 : Same as Figure 23 , for Dashcam Count..
Figure 27 : Same as Figure 23 , for Multi-View Seq..
Figure 28 : Same as Figure 23 , for Street-View Ord..
Figure 29 : Same as Figure 23 , for Surround Loc..
Figure 30 : Same as Figure 23 , for Traffic Causation.