Post-hoc attribution is widely used to identify important model inputs, but evaluating whether these attributions identify truly informative features is difficult because real datasets rarely provide reliable ground truth. We introduce a multimodal benchmark containing symbolic tabular, vision, and audio tasks with known informative feature sets. In the main benchmark conditions, informative variables, spatial regions, or channels occupy fixed input coordinates while the remaining inputs provide irrelevant or distracting information. We use the benchmark to study how attribution quality depends on predictive performance, irrelevant-feature burden and structure, model configuration, and attribution mechanism, while separately measuring identification, stability, and computational cost. In most symbolic-data sweeps, informative-set identification co-varies with predictive performance, while the relative ordering of attribution methods remains largely stable. Across modalities, no method dominates all architectures and conditions, and greater attribution complexity does not consistently improve identification. Simple gradient attribution is often competitive at lower computational cost, while the vision and audio results show that architecture and the form of irrelevant content materially affect attribution quality.
Figures & tables
Figure 1 : Benchmark motivation. Attribution scores are evaluated against a known informative feature set while irrelevant inputs and model conditions are varied.
Table 1 : Captum implementations and attribution parameters.
Figure 2 : Vision benchmark conditions. A class-bearing CIFAR-10 foreground is placed at a fixed (FP) or random (RP) position over a Gaussian random background (RB) or a structured but task-irrelevant Flower102 background (SB).
Figure 3 : Audio sources: (a) the predictive speech command, (b) Gaussian input, and (c) a structured but task-irrelevant environmental recording.
Data
Task
Predictive performance
Informative-set identification
Ground-truth informative set
Task-irrelevant inputs
Symbolic
Regression
NPS
RRA
Variables in the generating formula
Additional independent variables
Vision
Classification
Accuracy
IoU
CIFAR-10 foreground region
Gaussian noise or Flower102 background
Audio
Classification
Accuracy
RMA
Speech Commands channel
Gaussian noise or rainforest-audio channels
Table 2 : Benchmark design, evaluation, and ground-truth annotation.
Figure 4 : Symbolic one-factor sweeps. Colors and markers identify the six methods: ∙ DL, ▲ FA, ■ IG, ⧫ KS, × LIME, and ⊕ SA. The ▼ denotes prediction NPS. NPS and RRA are normalized to [0,1] , with larger values indicating better performance.
Figure 5 : Complementary symbolic-data evaluation. Convergence is the area under the RRA curve over 300 training epochs; consistency compares each sample’s top-ranked features with the dataset-average attribution. Both are normalized to SA. Memory and runtime are measured at test time. KS and LIME runtime values are scaled by (1/100) for visualization.
Panel A: Random background, fixed position (RBFP)
Architecture
Acc. (%)
SA
DL
IG
FA
KS
DLS
GB
LIME
ConvNeXt-Tiny ( Liu et al., 2022 )
85.11
0.829
0.666
0.554
0.860
0.021
0.666
0.656
0.021
DenseNet-121 ( Huang et al., 2017 )
83.38
0.738
0.566
0.601
0.748
0.022
0.674
0.639
0.022
EfficientNet-V2 ( Tan and Le, 2021 )
85.07
0.791
0.611
0.636
0.866
0.021
0.611
0.549
0.021
ResNet-18 ( He et al., 2016 )
77.44
0.678
0.530
0.551
0.677
0.020
0.532
0.450
0.020
ResNet-50 ( He et al., 2016 )
83.23
0.656
0.326
0.457
0.682
0.022
0.513
0.585
0.022
Table 3 : Vision accuracy and IoU under fixed-position random-background (RBFP) and structured-background (SBFP) conditions.
Panel A: Random background, fixed channel (RBFP)
Architecture
Acc. (%)
SA
DL
IG
FA
KS
DLS
GB
LIME
Conformer ( Gulati et al., 2020 )
20.97
0.310
0.015
0.025
0.225
0.277
0.025
0.100
0.320
LSTM ( Sak et al., 2014 )
73.77
0.684
0.590
0.260
0.515
0.431
0.248
0.697
0.564
RNN ( Medsker and Jain, 2001 )
67.00
0.653
0.585
0.275
0.447
0.427
0.440
0.814
0.548
TCN ( Kalchbrenner et al., 2016 )
80.96
0.641
0.297
0.258
0.539
0.139
0.185
0.663
0.325
Transformer ( Vaswani et al., 2017 )
23.25
0.460
0.236
0.138
0.335
0.170
0.110
0.440
0.305
Table 4 : Audio accuracy and RMA under fixed-channel Gaussian (RBFP) and structured task-irrelevant (SBFP) backgrounds.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
ID
∣S⋆∣
Generating function fj
1
1
a
2
1
a2
3
1
a2+12−1
4
1
sin(a)
5
1
exp(a)−1.5
6
1
2log(a2+1)−1
Appendix
Table 5 : Symbolic generating functions and informative-set sizes. Variables appearing in a formula are informative; all additional coordinates are task-irrelevant.
Carnegie Mellon University 1Department of Statistics & Data Science · Columbia University 2Department of Statistics · Columbia University 3Electrical Engineering Department +2