cs.LGOct 1, 2026

Invent a Dataset: Measuring dataset generation abilities with zero seed

Authors: Shivalika Singh, Andrija Djurisic, Gbemileke Onilude, Sudip Roy, Sara Hooker

Organizations: Adaption

Abstract

Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Autodata: An agentic data scientist to create high quality synthetic data

    Jun 24, 2026Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu +12Data Science AgentsSynthetic Data

  2. Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

    Apr 20, 2026Prasoon Goyal, Sattvik Sahai, Michael Johnston +14Large Language Model GenerationAdversarial Robustness

  3. ANDES: Agent Native Data Evolving Synthesis Tool for Autonomous Instruction Alignment

    May 31, 2026Zhengyang Zhao, Shengjie Ye, Lu Ma +3Data Science AgentsSynthetic Data