Scalable Attribution and Control of Model Behavior During Training
Organizations: Rice University
Abstract
Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguish. We address this ambiguity through mutual information, accounting for interference within the batch by quantifying how much the combined behavioral change reveals about each example's contribution. We show that this mutual information is a logarithmic function of Behavioral Gradient Uniqueness (BGU). BGU gives the information measure its geometric interpretation. Our Batch-Space Ghost (BS-Ghost) algorithm makes these scores practical inside the training loop through shared computation in batch space, without storing model-sized example gradients. On a complete 1,000-example Qwen2.5-7B-Instruct workload, our BS-Ghost implementation adds 27 seconds (8.0%) to 5.5 minutes of ordinary training. Removal and retraining demonstrate that BGU identifies data that causally shapes final behavior. At each training step, signed information identifies which examples strengthen or weaken the target behavior, explaining how behavior develops during training. Signed information also enables cheap intervention during training: it predicts how changing example weights will affect behavior in the next update. We then use these predictions to choose weights that steer behavior toward a desired target. This makes our framework a practical foundation for scalable oversight and verification of training pipelines and processes, helping evaluators assess model alignment, understand how it develops during training, and guide interventions that shape ongoing learning.
Figures & tables
| Run or method | Time | Timing Scope |
| Ordinary training | 331.90 s | same 1,000 examples and optimizer |
| Training with BS-Ghost | 358.51 s | every example scored during training |
| Added by BS-Ghost | 26.61 s | 8.02% increase; additional training time |
| Vector Filter ( Kowal et al., 2026 ) | 57 s | post-training pass |
| Projection Difference ( Chen et al., 2025 ) | 142 s | post-training pass |
| Concept Influence † ( Kowal et al., 2026 ) | 1,170 s | post-training pass |
| AUPRC | |||||
|---|---|---|---|---|---|
| Method | Execution | ToxicChat | XSTest | JailbreakBench | Post-training scoring (s) |
| LESS Xia et al. ( 2024 ) | Published | 0.388 | 0.724 | 1.000 | 5,400 |
| Reference V100 | 0.2710 | 0.7105 | 1.0000 | 4,704–4,739 | |
| BS-Ghost post-training | 0.2710 | 0.7105 | 1.0000 | 2,032–2,062 | |
| Grad-Dot Pruthi et al. ( 2020 ) | Published | 0.084 | 0.483 | 0.999 | 1,800 |
| Reference V100 | 0.0372 | 0.1752 | 0.9089 | 2,110–2,287 | |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Removed | Evil run 1 | Evil run 2 | Evil run 3 | Coherence mean | NLL |
|---|---|---|---|---|---|---|
| Opinions | 0% | 27.15 | 20.37 | 21.97 | 88.24 | 0 |
| 10% | 17.45 | 18.61 | 22.46 | 89.35 | .0089 | |
| 25% | 9.40 | 12.97 | 11.34 | 92.60 | .0276 | |
| 50% | 1.96 | 4.45 | 3.99 | 95.93 | .0711 | |
| 75% | .85 | 1.23 | 2.37 | 96.06 | .1302 | |
| 90% | 1.47 | 2.13 | 1.79 | 96.00 | .2362 |
| Training | Final persona | Boundaries above 30 | Exposure above 30 |
|---|---|---|---|
| Ordinary training | 43.03 | 555 | 22,923.5 |
| Response feedback | 10.29 | 154 | 6,100.1 |
| Signed-information feedback | 10.63 | 160 | 6,107.1 |
| Matched random | 35.67 | 555 | 21,404.6 |