cs.LGOct 13, 2025

Auditing Information Disclosure During Large-Scale Gradient-Based Training via Gradient Uniqueness

Authors: Sleem Abdelghafar, Maryam Aliakbarpour, Christopher Jermaine

Organizations: Rice University Houston, USA

Abstract

Auditing information disclosure across every datapoint during the training of LLMs is challenging. We propose a principled, attack-agnostic approach that uses mutual information to measure what the final model reveals about a datapoint's training membership. We show that this final-model disclosure is upper bounded by the sum of per-iteration gradient disclosures and that, under a reasonable set of assumptions, these gradient disclosures increase with Gradient Uniqueness (GNQ), which measures how distinguishable a datapoint's gradient is relative to other gradients in the batch. While naively computing GNQ requires forming and inverting a P×PP\times P matrix for every datapoint (for a model with PP parameters), we introduce Batch-Space Ghost (BS-Ghost). This efficient algorithm performs all computations in a much smaller batch space and uses ghost kernels to compute GNQ "in-run" for every datapoint in the training corpus, with minimal computational and memory overhead. Our experiments show the following: (i) GNQ predicts MIA vulnerability without the need for shadow models. (ii) Beyond membership disclosure, GNQ predicts the success of reconstruction attacks. (iii) GNQ-guided removal and retraining identify datapoints that causally contribute to disclosure. (iv) GNQ outperforms counterfactual memorization in text extraction and common-knowledge discrimination without the need for additional model training. (v) For data attribution, GNQ-guided filtering reduces emergent misalignment in Qwen2.5-7B more than baselines. Further, GNQ attributes 1000 datapoints in 17 seconds---roughly 70×70\times faster than the baselines. (vi) GNQ explains how training choices affect training-set disclosure and captures how per-datapoint disclosure emerges during training.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients

    Jun 23, 2026Zhihao Zhu, Hongyi Tang, Yi Yang +1Training DataGeneration Provenance

  2. Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training

    Jul 2, 2025Ismail Labiad, Mathurin Videau, Matthieu Kowalski +4Privacy-Preserving Machine LearningBlack-Box Optimization