Training data attribution (TDA) estimates the contribution of individual training examples to model outputs. Most scalable TDA methods rely on per-example gradients, whose computation and use at LLM scale pose challenges in efficiency, compatibility, and extensibility. We introduce dattri-LLM, a TDA library that makes gradient-based attribution more practical at scale. For efficiency, dattri-LLM uses compact gradient representations and dynamically routes gradient operations based on a cost model. For compatibility, its capture mechanism collects per-example gradients from existing training loops that call backward(), without requiring changes to the loop or its configuration. This includes distributed training with DDP and FSDP and pipelines built with HuggingFace Transformers, TRL, and OLMo. For extensibility, dattri-LLM exposes reusable gradient operations and training-time callbacks for implementing attribution methods and applications. These interfaces support a variety of attribution methods, including gradient similarity, curvature-based influence, and trajectory-based methods, as well as applications that act on gradients during training, such as online data selection. On the same hardware and workload, dattri-LLM achieves 3.2x the throughput of the fastest competing library on average, scales multiple attribution methods to 110B-parameter models across four H200 GPUs, and offers superior attribution fidelity-cost trade-offs across a range of models with different model families and scales. The source code of dattri-LLM is available at https://github.com/TRAIS-Lab/dattri-llm.
Figures & tables
Figure 1: Architecture overview of dattri-LLM . From bottom to top, the stack comprises external dependencies (gray), per-example gradient representations and operations (green), shared attribution components, and the entry API. Solid blue boxes denote classes that users instantiate and call; orange denotes configuration. Dashed boxes denote abstract base classes that serve as extension points: subclassing Attributor adds an attribution method, while subclassing HookManagerCallback adds an application that acts during gradient capture.
dattri-LLM
LogIX
Kronfluence
Bergson
Efficiency
Dimension reduction
✓
✓
✗
✓
Training-gradient reuse
✓
✓
✗
✓
Cost-based operation routing
✓
✗
✓
✗
Full gradients built on demand
✓
✗
∼
✗
Compatibility
Table 1: Comparison of recent libraries for LLM-scale TDA. ✓/✗ indicates supported / not supported; ∼ indicates partial support: Kronfluence uses a cost model, but it always builds full gradients on the query side and optionally builds training-side gradients on demand. Bergson integrates with a wrapped HuggingFace Trainer through callbacks that record gradients without acting on the training state. LogIX supports DDP but not FSDP. Kronfluence and Bergson provide interfaces for customizing specific steps of attribution, such as gradient collection or preconditioning, but not for defining a new end-to-end attribution method.
Figure 2: Throughput and memory scaling on the Qwen family. Top: throughput on four H200s at each configuration’s largest batch that completes within the benchmark limits. Crosses indicate configurations reported as out of memory. Bottom: GPU memory at batch size one. Filled markers denote single-H200 runs; hollow markers denote four-H200 sharded runs, with reported aggregate memory estimates. Dashed lines indicate aggregate device capacities, and crosses mark the total capacity at which an out-of-memory failure occurs.
Figure 3: Attribution fidelity versus throughput. Higher and farther right is better. Fidelity is measured as the mean query-wise Spearman correlation with TSLOO over 64 validation queries. Throughput is 64 divided by the attribution time in seconds on one B200. Dashed lines show the empirical Pareto frontier for each library. Configurations exceeding the resource budget are omitted.
Figure 4: Attribution-guided data selection under (a) full-parameter and (b) LoRA fine-tuning (SAMSum test perplexity, lower is better; mean over 3 seeds, band ±1 std). Both selection variants improve on training with all data, with per-layer selection best in both regimes.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Route
Arithmetic operations
Representation space
Factorized
O(B1B2T2(Ni+No))
O((B1+B2)T(Ni+No))
Materialized
O((B1+B2)TNiNo+B1B2NiNo)
O((B1+B2)NiNo)
Appendix
Table 2: Complexity of computing a cross-Gram matrix between two batches of per-example gradients. Space denotes the storage required for the gradient representations and excludes the O(B1B2) output matrix common to both routes. Under factor projection (Eq. 2 ), Ni and No are replaced by the projection dimensions ka and kg .
Scale
Checkpoint
Params
L
d
dff
Heads
Hooked
0.5 B
Qwen2.5-0.5B
0.49
24
896
4864
14/2
168
1.5 B
Qwen2.5-1.5B
1.54
28
1536
8960
12/2
196
3 B
Qwen2.5-3B
3.09
36
2048
11008
16/2
252
7 B
Qwen2.5-7B
7.62
28
3584
18944
28/4
196
14 B
Qwen2.5-14B
14.8
48
5120
13824
40/8
336
32 B
Qwen2.5-32B
32.5
64
5120
27648
40/8
448
Appendix
Table 3: Qwen models used in the scaling experiments: parameters (B), transformer blocks L , model width d , MLP width dff , attention heads (query / key-value), and hooked linear layers ( 7L ). The Qwen family provides a smooth progression in model size from 0.5 B to 110 B.
Method
Library
0.5 B
1.5 B
3 B
7 B
14 B
32 B
72 B
110 B
GradDot
dattri-LLM
613 (128)
299 (64)
186 (32)
110 (32)
49.9 (16)
22.0 (8)
8.6 (8)
6.9 (4)
Bergson
68.8 (64)
57.9 (64)
43.2 (32)
39.8 (32)
24.6 (16)
14.3 (8)
6.1 (8)
4.9 (4)
Kronfluence
177 (64)
111 (32)
64.4 (16)
41.5 (16)
—
—
—
—
LogIX
65.8 (128)
46.5 (64)
34.8 (32)
34.7 (32)
17.2 (16)
10.4 (8)
—
—
K-FAC
dattri-LLM
375 (128)
174 (64)
120 (32)
92.2 (32)
40.3 (16)
16.8 (8)
7.4 (8)
5.9 (4)
Bergson
2.8 (32)
1.9 (32)
1.2 (16)
0.8 (8)
—
—
—
—
Appendix
Table 4: Throughput (training sequences/s) on four H200s, with the per-device batch size in parentheses. The workload contains 512 training sequences and one query through 32 B, and four queries at 72 B and 110 B. A dash denotes a configuration reported as out of memory.
Method
Library
0.5 B
1.5 B
3 B
7 B
14 B
32 B
72 B
110 B
GradDot
dattri-LLM
117.6
93.0
67.2
88.4
72.0
79.0
131.0
127.8
Bergson
80.9
119.0
86.5
104.6
87.7
91.2
139.8
135.0
Kronfluence
101.7
95.3
80.7
119.7
—
—
—
—
LogIX
117.4
94.6
70.9
97.4
91.0
124.4
—
—
K-FAC
dattri-LLM
117.6
93.0
67.2
88.4
72.0
80.0
132.1
127.8
Bergson
95.4
134.4
112.4
132.5
—
—
—
—
Appendix
Table 5: Peak per-device GPU memory (GB) at the batch sizes used in Table 4 . Batch sizes vary across methods, libraries, and scales; these values document the throughput configurations and do not constitute a fixed-batch memory comparison.
Method
Scale
dattri-LLM
Bergson
LogIX
Kronfluence
GradDot
0.5 B
2
2.1
2
3.7
1.5 B
4.8
4.8
4.9
10.2
3 B
8.7
8.5
8.6
19.9
7 B
18.8
17.3
17.8
43.5
14 B
33.8
32.2
33.2
85
32 B
72.1
69.9
69.8
—
Appendix
Table 6: GPU memory usage (GB) at batch size one, with 64 training sequences and one query. F denotes a completed sharded run on four H200s, with memory reported as an estimated aggregate across all four devices. All other completed runs use one H200. A dash indicates an out-of-memory failure under the evaluation protocol.
Method
Model
Limiting resource
MAGIC (Bergson)
Qwen2.5- 1.5 B
GPU memory
AdamW-influence (full, dattri-LLM )
Qwen2.5- 1.5 B
GPU memory
MAGIC
OLMo-3- 7 B
GPU memory
AdamW-influence (full)
OLMo-3- 7 B
GPU memory
SOURCE (Bergson)
OLMo-3- 7 B
Local disk
EK-FAC (Bergson)
OLMo-3- 7 B
GPU memory
Appendix
Table 7: Resource limits in the fidelity benchmark under the budget of one B200, 256 GB of host memory, and 1 TB of local disk.
GPT-2
Qwen2.5- 1.5 B
OLMo-3- 7 B
Library
Method
ρ
tattr
GB
ρ
tattr
GB
ρ
tattr
GB
dattri-LLM
AdamW-inf. ( k=512 )
.680
10
8
.526
51
47
.437
73
136
dattri-LLM
AdamW-inf. ( k=8192 )
.762
28
8
.639
103
49
.633
170
138
dattri-LLM
AdamW-inf. (full)
.845
1137
25
–
–
–
–
–
–
dattri-LLM
EK-FAC ( ka=kg=64 )
.699
15
8
.420
50
47
.303
126
40
dattri-LLM
EK-FAC (full)
.664
17
30
.434
447
96
.402
3444
139
Appendix
Table 8: Complete fidelity results on one B200: Spearman correlation with TSLOO ( ρ ), attribution time of the 64 queries in seconds ( tattr ), and peak GPU memory in GB. “–”: infeasible under the budget.
dattri-LLM
LogIX
Bergson
Kronfluence
Method
Proj.
tattr
MGPU
tattr
MGPU
tattr
MGPU
tattr
MGPU
One query ( ntest=1 )
GradDot
ka=kg=64
41
7
68
7
39
8
n/a
n/a
full
45
11
oom
—
53
13
50
11
K-FAC
ka=kg=64
54
10
89
7
180
20
n/a
n/a
full
134
19
oom
—
194
20
201
12
Appendix
Table 9: Cross-library efficiency under matched batch sizes. Results use Pythia- 410 M on one A40 ( 46 GB), with 1024 WikiText-103 training sequences and either one or 16 queries. tattr : attribution wall-clock time (s), excluding model construction and data loading; MGPU : peak GPU memory (GB) over the same phases. The lowest runtime in each row is bold. Completed configurations use training batch size 8 . oom indicates an out-of-memory failure at every tested batch size down to one; n/a indicates an unsupported combination.
Figure 7: Adaptive routing versus fixed gradient representations. Mean attribution time per step (left) and peak GPU memory (right) for full-dimensional GradDot on Pythia- 410 M as sequence length increases, using batch size 8 , one query, and fp32 on one A40. The cost-model heuristic selects between factorized and materialized execution; fixed-route ablations use the same implementation. Adaptive routing closely tracks the faster fixed configuration while maintaining the lowest or tied peak memory at every evaluated length. Memory includes all GPU allocations.
Figure 8: maskray (train row 308). Target: “The plain maskray is a species of stingray that feeds on caridean → shrimp ”. Under GradDot the largest positions are the subword ts of hunts , mask , and the two occurrences of of , with shrimp itself at rank 6 . Both preconditioned methods concentrate on of caridean , the phrase the target completes.
Figure 9: jordan (train row 1724). Target: “Michael Jordan attended Emsley A. Laney High School in → Wilmington ”. The paragraph states the fact. GradDot ranks Wil second, behind the function word in , and shades the whole paragraph; K-FAC and EK-FAC put Wil first, followed by in and the remaining subwords of the city name, and leave the text after the school name essentially blank.
Figure 10: slayer (train row 752). Target: “South of Heaven is a studio album by the American thrash metal band → Slayer ” (tokenized Sl ayer ). The answer’s first subword ranks first under all three attributors, but its lead over the runner-up grows from 3.6× ( GradDot ) to 7.3× (EK-FAC) and 35.7× (K-FAC). The paragraph discusses the album’s tempo rather than restating the band’s name, and the sequence-level GradDot score is in fact negative ( −82 ), whereas both preconditioned scores are positive: preconditioning concentrates positive attribution on the answer token, while the unpreconditioned sequence-level score is negative.
Method
Projection
ntest=1
ntest=16
GradDot
ka=kg=64
1.19×
1.18×
GradDot
full
1.16×
1.09×
K-FAC
ka=kg=64
1.14×
1.11×
K-FAC
full
1.10×
1.07×
EK-FAC
ka=kg=64
1.10×
1.14×
EK-FAC
full
1.08×
1.08×
Appendix
Table 10: Speedup of weight-gradient-free capture over standard capture (standard time / weight-gradient-free time) for dattri-LLM on Pythia- 410 M, one L40S (48 GB), fp32, with the workload of Table 9 . Each entry is the median over five paired repetitions.
Figure 11: Per-step runtime of online data selection relative to standard training, under full-parameter fine-tuning.