Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging. Existing attribution methods often compress factor-specific effects into scalar responses, making distinct internal changes indistinguishable. This is particularly limiting for diffusion models, where semantic factors emerge through evolving representation dynamics during denoising. We therefore reformulate diffusion data attribution as attributing factor-induced internal response trajectories. In this paper, we propose a novel Concept Attribution method through Dynamic Trajectories(CADT). We argue that attribution should therefore ask not only \emph{which} examples matter, but also \emph{how} their influence unfolds during generation. Specifically, we construct matched counterfactual pairs at identical noisy states to isolate factor-specific representation displacements, and model their directional and magnitude evolution across denoising as dynamic attribution signatures. For each training example and generated query, CADT extracts stage-wise feature vectors and integrates them along the denoising process to form a trajectory descriptor. Applying the same construction across the training set yields a bank of factor-specific trajectory descriptors. The covariance statistics of this bank are then used to construct . CADT uses this covariance-aware positive-semidefinite kernel to calibrate the query and training representations, and compares the calibrated query trajectory with each training trajectory to produce the final training-sample attribution scores. Experiments on multiple public datasets show consistent improvements over existing diffusion attribution baselines across hierarchical, compositional, and style attribution.
Figures & tables
Figure 1: The overview of our proposed CADT . Throughout the experiments, the term "style" denotes either color or image style.
Figure 2: Synthetic setup for shape and color attribution.
Shape
Color
Avg.
Method
P@1
P@5
P@10
P@1
P@5
P@10
P@1
P@5
P@10
TRAK
71.20
41.10
33.64
26.20
47.54
50.96
48.70
44.32
42.30
D-TRAK
0.30
0.68
1.00
99.70
99.32
99.00
50.00
50.00
50.00
DAS
65.30
43.44
37.27
34.70
53.70
50.52
50.00
48.57
43.90
CADT
91.20
91.36
92.01
90.80
93.88
93.65
91.00
92.62
92.83
Table 1: Synthetic held-out OOD precision (%).
Figure 3: Signed alignment of the query with positive and negative training examples.
Figure 5
Method
Mid-level
Fine class
D-TRAK
45.30
18.54
TRAK
41.35
15.30
DAS
11.95
1.62
Journey-TRAK
10.76
1.95
CADT
50.59
21.68
Table 2: CIFAR-100 P@5 (%).
Figure 6: Label-corruption retrieval for the same query, showing top-five training examples. Red: corrupted animals; green: genuine plants. Top: D-TRAK; bottom: CADT .
Method
Avg.
Flipped animal → plant
Genuine plant
TRAK
56.22 / 52.00 / 50.14 / 44.27
23.24 / 21.95 / 20.00 / 16.40
89.19 / 82.05 / 80.27 / 72.14
D-TRAK
63.78 / 59.89 / 57.68 / 49.12
40.54 / 35.68 / 33.51 / 26.30
87.03 / 84.11 / 81.84 / 71.94
DAS
24.05 / 21.46 / 19.89 / 22.28
1.08 / 0.76 / 1.03 / 4.17
47.03 / 42.16 / 38.76 / 40.39
Journey-TRAK
1.35 / 6.59 / 22.57 / 21.75
1.08 / 1.51 / 13.51 / 9.01
1.62 / 11.68 / 31.62 / 34.49
CADT
67.57 / 66.86 / 66.51 / 64.48
37.84 / 38.27 / 37.51 / 34.42
97.30 / 95.46 / 95.51 / 94.53
Table 3: CIFAR-100 20% label-flip source recovery (%). Values within each cell are P@ 1/5/10/50 .
Table 6: Ablation of descriptor components and temporal support on CIFAR-100 at P@5.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Score
Macro
Flipped animal → plant
Genuine plant
CADT ( Kλ )
67.57 / 66.86 / 66.51 / 64.48
37.84 / 38.27 / 37.51 / 34.42
97.30 / 95.46 / 95.51 / 94.53
Normalized variant ( Sλ )
70.54 / 68.27 / 67.35 / 65.32
44.32 / 40.86 / 39.78 / 36.08
96.76 / 95.68 / 94.92 / 94.56
Ordinary cosine
74.59 / 73.41 / 72.19 / 70.58
53.51 / 52.54 / 50.70 / 48.04
95.68 / 94.27 / 93.68 / 93.11
Appendix
Table A1: Exploratory score-design comparison for CIFAR-100 label-flip recovery. Values are P@ 1/5/10/50 . All rows use the identical full late-support descriptor, query bank, and candidate order; only the scoring function changes.
Figure A1: Fine-grained direction and energy patterns in the selected low-noise region. Representative positive and negative pairs are selected according to their separation under Kλ .
Figure A2: Matched-minus-mismatched direction, energy, and Kλ cell margins in the fine-grained 200→100 analysis.
Target
Direction– Kλρ
Energy– Kλρ
Selected layer: gap [95% CI] ( 10−8 )
Selected window: gap [95% CI] ( 10−6 )
Shape
0.728
−0.109
d1: 2.01[1.36,2.65]
300→100 : 0.871[0.852,0.890]
Color
0.972
0.025
d0: 9.80[9.69,9.91]
300→100 : 1.097[1.078,1.117]
Appendix
Table A2: Full-range layer–time analysis under Kλ . Correlations summarize the correspondence of direction and energy with Kλ , while the final columns report the selected layer and temporal window together with their separation from the strongest alternative.
Figure A3: Examples from the fixed generated-query bank for ArtBench5. The ArtBench2 setting uses a subset of the styles included in ArtBench5, and is therefore covered by the examples shown here. Query validity is judged once and the same mask is shared by every method.
Figure A4: All 40 ArtBench5 cat–style composition generations before the shared validity mask (five styles, eight seeds each). This grid documents query coverage; only queries judged to contain both the requested object and style enter Table 5 .
Figure A5: Representative style retrieval results produced by CADT . (a) ArtBench2, post-impressionism. (b) ArtBench5, baroque. Blue denotes the generated query and green denotes style-relevant retrieved training examples.
Figure A6: Changes in denoising error after style-specific and size-matched random deletion. Rows correspond to held-out queries and columns to denoising timesteps. The right panels show the style-minus-random contrast, revealing stronger and temporally structured changes for queries associated with the removed training style.
Figure A7: CADT deletion changes distribution of retrained ArtBench2 models.
Figure A8: Scalar-matched examples for fine-class attribution on CIFAR-100. (a) A generated butterfly query. (b) and (c) Two butterfly training examples with nearly equal estimated Concept-TRAK scores. The pair is selected by score proximity under a fixed utility-time sample.
Figure A9: Query-time by training-time contributions to the Concept-TRAK score. The two maps have nearly equal total sums but exhibit different temporal structure.
Figure A10: Query-time attribution profiles obtained by aggregating the contribution maps over training timesteps. Similar scalar scores can arise from different distributions of attribution across denoising.
Figure A11: Predicted clean images along the reconstruction trajectories of two scalar-matched training examples. The examples exhibit different denoising progressions despite their similar Concept-TRAK scores.
Figure A12: Layer–time magnitudes of the fine-class versus null-conditioned response for the two scalar-matched examples. Responses are evaluated at identical noisy states within each example.
Training data attribution for diffusion models aims to identify the training samples that influence a generated instance, but existing methods either require costly per-sample gradient computation or query-specific model optimization. Moreover, most methods attribute changes in a proxy loss rather than changes in the actual model's generative behavior. We address these limitations by formulating attribution directly with a local score discrepancy measure, which applies to any diffusion variant (including DDPM, EDM, and flow matching), and by showing that such measure can be estimated without retraining, as a preconditioned gradient similarity. We instantiate this estimator as Training-data Influence via score Discrepancy (TID), which uses Kronecker-factored curvature to avoid random projections and per-sample gradient storage. We then distill TID into TIDE, a forward-only student trained online to reproduce the teacher's rankings from the diffusion model's internal activations. Under counterfactual evaluation on CIFAR-10, ArtBench-10, and MS-COCO, TID matches or outperforms state-of-the-art approaches, while TIDE retains most of TID's accuracy at four to five orders of magnitude lower per-query cost, attributing generated samples in milliseconds and faster than the generation itself.
Shixuan Liu, Joan Serrà, Kin Wai Cheuk +6
University of Illinois at Urbana-Champaign · Sony AI · University of Texas at Austin +1
Attributing a generated image to its source diffusion model is a fundamental challenge in provenance verification and intellectual property protection. This problem is particularly difficult because diffusion models trained on different datasets can converge to similar score functions and thus similar output distributions, making the generated images themselves unreliable as attribution evidence. Existing non-invasive methods either fail on architecturally similar variants or rely on signals that vanish when models share the same autoencoder. We propose Spectral Denoising Signatures (SDS), a non-invasive attribution method that identifies the source model by fingerprinting each candidate model's denoising behavior. Our key insight is that a model's denoising score function exhibits a distinctive spectral geometry, reflected in how it redistributes energy across spatial frequency bands during denoising. By probing this behavior with frequency-controlled perturbations, SDS extracts a stable signature that is intrinsic to the model, requiring only standard forward passes with no inversion, optimization, or generation-time enrollment. Our results demonstrate that SDS achieves approximately 99.9% accuracy across eight diverse diffusion models and 96.2% under cross-domain prompt shift, outperforming non-invasive baselines across variations in training data, architecture, and training procedure, establishing spectral geometry as a principled and practical basis for diffusion model attribution. Code is available at: https://github.com/Pragati-Meshram/SGS
Pragati Shuddhodhan Meshram, Varun Chandrasekaran
Department of Electrical and Computer Engineering University of Illinois Urbana-Champaign
Training data attribution (TDA) should enable generative model interpretability and foster a variety of related downstream tasks. Nonetheless, current TDA approaches lack reliability and robustness, preventing their adoption in real-world setups. In this paper, we take a decisive step towards more reliable and robust TDA for diffusion models. We propose to perform TDA with mirrored unlearning and noise-consistent skew (MUCS). The idea is to fine-tune a second model with bounded mirrored gradient ascent, and to measure the normalized skew of this model with respect to the original one using consistent noise samples. We show that, while being conceptually simple and generic, MUCS systematically outperforms existing methods on three different datasets by a large margin. We additionally study the effect that core design choices have on final performance, and analyze novel aspects regarding the overlap of influential instances across generated items and the potential of ensembling TDA approaches. We believe that our findings may have broader implications for more general unlearning setups, as well as for tasks requiring the comparison of diffusion losses.