cs.LGSep 26, 2026

Traceprop: Training Data Attribution in a Single Pass

Authors: Amit Nautiyal

Organizations: Independent Researcher

Abstract

Data attribution tools used in practice, TRAK, LoGRA/LogIX, and EK-FAC, run after training: they recompute per-sample gradients in a separate pass over the training set, and curvature-aware variants need a further pass to estimate covariance. Traceprop records projected per-sample gradients and K-FAC covariance statistics inside the training backward pass itself. On Pythia-1B LoRA fine-tunes on an A100, this adds 2.6% to training time, against 10.7% for LogIX's best inline configuration, random-init (4.1x lower, one-sided Mann-Whitney p = 9e-5, n = 10 per arm); at Pythia-6.9B the numbers are 11.3% and 27.0% (p = 0.004). LogIX's default configuration, PCA-init, is both slower and lower quality than random-init, so we compare against random-init throughout. On a small transformer where LDS is measurable, Traceprop matches LogIX's best configuration: pooled LDS difference +0.0022, 95% CI [-0.0038, +0.0075]. Two things do not work: at Pythia-160M/SST-2, every method is indistinguishable from a low noise ceiling, and on planted-backdoor and mislabel detection, gradient attribution does not beat gradient norm or representation similarity. The contribution here is systems, not a new estimator: LogIX's attribution quality, in one pass instead of two.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale

    Sep 30, 2026Shixuan Liu, Tongli Zhou, Junwei Deng +2Large Language Model TrainingLarge Language Models(Llms

  2. STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

    Jun 3, 2026Rishit Dagli, Abir Harrasse, Luke Zhang +4Training Data

  3. How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines

    May 12, 2026Junwei Deng, Pingbang Hu, Suliang Jin +4Alignment-Aware SelectionLog-Ratio-Based Gaussian Trust Weight