cs.CLSep 30, 2026

Synthetic Data Characterization via Training Dynamics

Authors: Irene Lago, Ana Ezquerro, David Vilares

Organizations: Universidade da Coruña, CITIC, Spain · Graz University of Technology, IML, Austria

Abstract

Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Effective Synthetic Data Curation Requires Group-Level Signals

    Sep 30, 2026Cathy Jiao, Chenyan XiongSynthetic DataData-Curation

  2. Training-Aware Target Coverage for Synthetic Data Selection

    Sep 30, 2026Yang Ba, Michelle V. Mancenido, Rong PanSynthetic DataLarge Language Model Fine-Tuning

  3. An Information-Theoretic Criterion for Efficient Data Synthesis

    May 11, 2026Hanyu Li, Zhengqi Sun, Xiaotie DengSynthetic DataInformation Bottleneck