End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84% of the gain from supplying all prerequisites.
Figures & tables
Figure 1 : An example CADET task ( left ) decomposed into 6 unit tasks formalized as an SCM ( middle ). Right : natural capability (NC) of each unit task ( top ), and their causal contribution ( N / S(1,∅) ; bottom ) to final task-6, where supplying unit task-2 alone raises final accuracy from 0.27 to 0.65.
Contrast
N-Score (Success drop from ⋯ )
S-Score (Failure recovery from ⋯ )
(1,0)
corrupting Xi into a wrong answer in oracle env.
fixing Xi into the correct answer in all-wrong env.
(1,∅)
leaving Xi unassisted in oracle env.
providing the correct Xi alone in natural env.
(∅,0)
injecting a wrong answer for Xi in natural env.
removing a wrong answer for Xi in all-wrong env.
(0,∅)
removing Xi ’s wrong answer in all-wrong env.
injecting a wrong Xi as a scaffold in natural env.
Table 1: Operational semantics of N-Score, S-Score across 4 ordered state contrasts (a,a′) .
Figure 2 : Overview of SCMs corresponding to 10 tasks in CADET.
Figure 3 : Natural(NC) and intrinsic(IC) capabilities on non-root unit tasks: (a) category-level NC and IC averaged over 12 models; (b–e) per-model NC and IC within each category.
Figure 6 : N/S(1,∅) across Pattern 1( a,b ) and Pattern 2 ( c,d ) tasks for top-4 SOTA ( a,c ) and remaining 8 Others ( b,d ) models ( sqrt(⋅) -scaled axes, similar in Fig 5 ).
Figure 6Figure 7
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Task Name
Modality
Category
# Instances
Answer Format
Music Note Counting
Image
Perception
114
Integer
Dice Sum
Image
Perception
130
Integer
Map Navigation
Image
Spatial
107
Integer
Maze
Image
Spatial
80
Multiple-Choice
Surround Localization
Multi-Image
Spatial
154
Integer
Multi-View Sequencing
Multi-Image
Temporal
154
Letter List
Appendix
Table 2 : Overview of the 10 overall evaluation tasks in CADET.
Task
#Units
#Roots
Depth
∣Pa(Y)∣
Category
#Questions
Has Markov
Music Note Counting
4
2
2
3
P/C
1,163
No
Dice Sum
4
2
2
3
P/C
1,919
No
Map Navigation
3
1
2
2
P/S/C
918
Yes
Maze
2
1
1
1
P/S
880
Yes
Surround Localization
6
3
2
5
P/S
4,928
No
Multi-View Sequencing
6
2
3
5
P/S/T
4,158
No
Appendix
Table 3 : Structure of the task-level SCM of each evaluation task. – on multistep reasoning is due to it start from a Markov chain structure.
Node
Unit Task
Cat.
Answer Format
#Q
Parents
Music Note Counting
X1
Object Identification
P
Multiple-choice
821
—
X2
Object Classification
P
Multiple-choice
114
X1
X3
Completeness Verification
P
Multiple-choice
114
—
Y
Object Counting
C
Integer
114
X1,X2,X3
Dice Sum
Appendix
Table 4 : Unit task specification of the 10 evaluation tasks in CADET. Y denotes the overall task outcome; ↻ marks a Markov chain unit task.
Figure 11 : Prompt template for our evaluation under active causal interventions ( Ai∈{1,0} ). Under the natural state ( Ai=∅ ), only the target question and last paragraph of output format instructions are prompted.
Figure 12 : LLM-as-a-judge prompt template.
Figure 13 : Comparing NC on root unit tasks against NC and IC on non-root unit tasks, per category, averaged over the 12 models. Cognitive is not shown as it has no root unit task. The non-root bar stacks IC − NC on top of NC, so its top is IC.
Figure 14 : Comparing NC on root unit tasks against NC and IC on non-root unit tasks for each of the 12 models. There are 16 root unit tasks and 30 non-root unit tasks. The non-root bar stacks IC − NC on top of NC, so its top is IC. Models are ordered by overall non-root NC.
Figure 15 : Comparing NC and RC across the four categories on non-root unit tasks, averaged over the 12 models. Each bar shows NC with the signed NC/RC difference stacked out of it, so the boundary of the shaded band is RC.
Figure 16 : Comparing NC and RC for each of the 12 models within each category on non-root unit tasks. Circles mark NC and triangles mark RC; the shaded band spans the two and is colored by the sign of RC − NC. The horizontal lines are the means over the 12 models.
#unit tasks
Root
Non-root
Category
root
non-root
NC
NC
IC
IC − NC
Perception
11
4
0.775 ± 0.077
0.590 ± 0.138
0.796 ± 0.112
+0.206
Spatial
3
9
0.627 ± 0.100
0.556 ± 0.077
0.671 ± 0.059
+0.115
Temporal
2
8
0.714 ± 0.119
0.460 ± 0.072
0.614 ± 0.087
+0.154
Cognitive
0
9
–
0.406 ± 0.092
0.721 ± 0.089
+0.316
Appendix
Table 6 : NC on root unit tasks and NC, IC on non-root unit tasks, per category, with the number of unit tasks in each group. Entries are the unweighted mean over unit tasks, averaged over the 12 models, ± the standard deviation across models.
Model
Root NC
Non-root NC
Non-root IC
NC(root) − NC(non-root)
NC(root) − IC(non-root)
GPT-5.4
0.781 ± 0.108
0.612 ± 0.180
0.791 ± 0.167
+0.169
− 0.010
Gem 3.5F
0.839 ± 0.099
0.595 ± 0.217
0.783 ± 0.172
+0.244
+0.056
Gem 3.1P
0.842 ± 0.090
0.577 ± 0.202
0.781 ± 0.173
+0.265
+0.062
Gem 3F
0.811 ± 0.119
0.528 ± 0.206
0.757 ± 0.180
+0.282
+0.053
Qwen27B
0.780 ± 0.132
0.515 ± 0.247
0.711 ± 0.219
+0.264
+0.069
Qwen397B
0.749 ± 0.144
0.487 ± 0.240
0.647 ± 0.226
+0.262
+0.101
Appendix
Table 7 : NC on root unit tasks and NC, IC on non-root unit tasks for each of the 12 models. Entries are the unweighted mean over unit tasks ± the standard deviation across unit tasks (16 root, 30 non-root). The last two columns are differences of the corresponding means.
Category
NC
IC
RC
IC − NC
RC − NC
Perception
0.590 ± 0.138
0.796 ± 0.112
0.446 ± 0.131
+0.206
− 0.144
Spatial
0.556 ± 0.077
0.671 ± 0.059
0.439 ± 0.102
+0.115
− 0.116
Temporal
0.460 ± 0.072
0.614 ± 0.087
0.414 ± 0.063
+0.154
− 0.046
Cognitive
0.406 ± 0.092
0.721 ± 0.089
0.445 ± 0.060
+0.316
+0.040
Appendix
Table 8 : NC, IC and RC per category on the 30 non-root unit tasks. Entries are the unweighted mean over unit tasks, averaged over the 12 models, ± the standard deviation across models.
Perception
Spatial
Temporal
Cognitive
Overall
Model
NC
IC
RC
NC
IC
RC
NC
IC
RC
NC
IC
RC
NC
IC
RC
GPT-5.4
0.636 ± 0.194
0.913 ± 0.087
0.507 ± 0.288
0.665 ± 0.202
0.747 ± 0.214
0.549 ± 0.213
0.560 ± 0.148
0.730 ± 0.181
0.537 ± 0.219
0.593 ± 0.192
0.835 ± 0.094
0.575 ± 0.274
0.612 ± 0.180
0.791 ± 0.167
0.548 ± 0.232
Gem 3.5F
0.722 ± 0.192
0.889 ± 0.042
0.599 ± 0.194
0.672 ± 0.231
0.754 ± 0.213
0.572 ± 0.252
0.564 ± 0.154
0.699 ± 0.202
0.481 ± 0.187
0.490 ± 0.232
0.839 ± 0.091
0.507 ± 0.219
0.595 ± 0.217
0.783 ± 0.172
0.532 ± 0.212
Gem 3.1P
0.761 ± 0.155
0.907 ± 0.027
0.678 ± 0.158
0.650 ± 0.217
0.740 ± 0.230
0.584 ± 0.207
0.515 ± 0.164
0.698 ± 0.175
0.450 ± 0.222
0.478 ± 0.174
0.839 ± 0.088
0.513 ± 0.194
0.577 ± 0.202
0.781 ± 0.173
0.540 ± 0.206
Gem 3F
0.664 ± 0.250
0.879 ± 0.072
0.439 ± 0.304
0.570 ± 0.268
0.710 ± 0.253
0.431 ± 0.240
0.524 ± 0.122
0.707 ± 0.163
0.422 ± 0.205
0.430 ± 0.154
0.796 ± 0.121
0.450 ± 0.223
0.528 ± 0.206
0.757 ± 0.180
0.435 ± 0.222
Qwen27B
0.682 ± 0.167
0.873 ± 0.119
0.418 ± 0.237
0.571 ± 0.308
0.663 ± 0.252
0.439 ± 0.249
0.515 ± 0.248
0.672 ± 0.210
0.466 ± 0.276
0.386 ± 0.159
0.722 ± 0.220
0.421 ± 0.213
0.515 ± 0.247
0.711 ± 0.219
0.438 ± 0.233
Appendix
Table 9 : NC, IC and RC per model and category on the 30 non-root unit tasks, with Overall aggregating all 30. Entries are the unweighted mean over unit tasks ± the standard deviation across unit tasks.
Figure 17 : Prerequisite N/S -Scores under state contrasts (a) (1,0) , (b) (0,∅) , and (c) (∅,0) , averaged across models. Dashed circles mark the highest- S prerequisite within each task. Axes and quadrant thresholds follow Figure 5 .
Figure 18 : Prerequisite N/S -Scores under (1,0) (a–d), (0,∅) (e–h), and (∅,0) (i–l) by task pattern and model group. Large markers denote group means; grey markers indicate individual models.
Figure 19 : Pairwise Spearman rank correlation between models on N -Scores (a–d) and S -Scores (e–h) across prerequisites under the four state contrasts. Lines separate SOTA from Others.
Figure 20 : Prerequisite N/S -Scores under the four state contrasts, averaged over evaluated models.
Figure 25
Figure 23 : Per-model N - and S -Scores for the parents of Map Nav. under the four state contrasts. The vertical line separates SOTA from Others.
Figure 24 : Same as Figure 23 , for Dice Sum.
Figure 25 : Same as Figure 23 , for Music Note.
Figure 26 : Same as Figure 23 , for Dashcam Count..
Figure 27 : Same as Figure 23 , for Multi-View Seq..
Figure 28 : Same as Figure 23 , for Street-View Ord..
Figure 29 : Same as Figure 23 , for Surround Loc..
Figure 30 : Same as Figure 23 , for Traffic Causation.
Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses visible, we propose Scaffolded Task Design (STaD) framework. STaD generates controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. Rather than inspecting failures individually, this approach enables systematic and scalable probing of model behavior by identifying the specific reasoning skill compositions they lack. Treating the LLM as a black box, our experiments on six models of varying sizes reveal multiple failure points in three reasoning benchmarks and highlight each model's unique and distinct skill gaps.
Sungeun An, Swanand Ravindra Kadhe, Shailja Thakur +2
Humans cannot always intuit what scenarios are most challenging to LLMs. Hoping to capture challenging edge cases, developers either design problems to be difficult for humans or curate extensive benchmarks. What if we could instead anticipate which scenarios a model will fail on? In this paper, we use an LLM's representational geometry to predict which concept combinations it will fail on. We attribute this compositional failure to interference between salient features. In tasks that require systematic composition - toy programmatic settings, multihop reasoning, multilingual factual recall - we find that when a pair of concepts is encoded near-orthogonally, the model reliably composes them. When their linear encodings are close, producing interference, the model fails to compose them. Our method reliably anticipates failure modes across different compositional tasks, without evaluating specific inputs. These results lay the groundwork to use representational geometry to identify high-risk examples, construct targeted stress tests, and provide a scalable foundation for active learning in real-world deployment.
Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3
Brown University · University of Southern California · Harvard University +1
Standard benchmarks report aggregate accuracy, but practitioners need to know which specific capabilities a model lacks. We introduce FailureScope, a behavioral-diagnosis method that clusters evaluation probes by their cross-model pass/fail patterns (leave-one-model-out, LOMO), and show it yields stable, interpretable failure taxonomies across three regimes usually studied separately: single-turn benchmarks, multi-turn dialogue, and adversarial agent attacks. On 2,664 single-turn tasks across 18 models, taxonomy-conditioned sampling reaches Kendall's tau = 0.81 at 50 tasks (versus 0.34 for random selection), and cross-model failure prediction reaches AUC 0.88. The same primitive recovers interpretable clusters on a 363-task multi-turn corpus and on 630 adversarial agent traces, where it exposes a meta-failure mode: a 73-100 percentage-point gap between LLM-judge ASR and real execution. Cluster cohesion remains strong across all three regimes, which we take as evidence that behavioral clustering is a portable diagnosis primitive that generalizes beyond any single benchmark. We release the pipeline, three annotated corpora, and the cross-regime taxonomies.