Training Data

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

13 new papers

A weekly snapshot of new work published in Training Data.

Inside this field

Focused directions

312 papers

Latest in Training Data

  1. DMuon: Efficient Distributed Muon Training with Near-Adam Overhead

    Jun 25, 2026Vincent Chen, Starrick Liu, Regis Cheng +8Muon OptimizerMuon

  2. Dataset Usage Inference without Shadow Models or Held-out Data

    Jun 24, 2026Wojciech Łapacz, Stanisław Pawlak, Jan Dubiński +2Training DataPrivate Data

  3. Less is More: Quality-Aware Training Data Selection for Scientific Summarization

    Jun 23, 2026Maria Nefeli Paraskevopoulou, Tatiana Passali, Grigorios TsoumakasSummarizationBiomedical Text

  4. BluTrain: A C++/CUDA Framework for AI Systems

    Jun 23, 2026Adhitya Charan, Adwaid Suresh, Anuj Kumar +19CudaTrainability

  5. MGI: Member vs Generated Inference

    Jun 22, 2026Bihe Zhao, Michel Meintz, Juangui Xu +2Membership InferenceGenerative Models

  6. ProCUA-SFT Technical Report

    Jun 15, 2026Jaehun Jung, Ximing Lu, Brandon Cui +11Computer-Use AgentsAgent Framework

  7. Piper: A Programmable Distributed Training System

    Jun 9, 2026Megan Frisella, Shubham Tiwari, Andy Ruan +5ParallelModel Training

  8. Massive Open-Vocabulary Keyword Spotting

    Jun 9, 2026Leonor Barreiros, Raul Monteiro, Afonso Mendes +1Open-VocabularyTraining Data

  9. Integral Field Unit Spectroscopy with One Fiber

    Jun 8, 2026Zehao Peng, Biprateep Dey, Chris J. Maddison +1SpectraCosmology

  10. Logit Distillation on Manifolds: Mapping by Learning

    May 30, 2026Yiru Yang, Junling Wang, Nishant Kumar Singh +2Model TrainingProgressive Distillation