Invent a Dataset: Measuring dataset generation abilities with zero seed
Organizations: Adaption
Abstract
Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.
Figures & tables
| Dataset | Unconstrained Query | Constrained Query |
| African QA | Dataset containing question-and-answer pairs covering diverse topics related to African research, literature, history, and current events. | Dataset containing multi-turn conversations between a teacher and a student on an exam , covering diverse topics related to African research, literature, history, and current events. Prompts must use OpenAI format while responses should be plain text. |
| Legal QA | Dataset containing question–answer pairs focusing on legal advice, factual explanations, and comparative analysis across diverse topics such as copyright law, government types, and liability. | Dataset containing question–answer pairs focusing on legal advice, factual explanations, and comparative analysis across diverse topics such as copyright law, government types, and liability. Questions should offer four choices; answers are limited to A, B, C, or D with no additional text. |
| Serbian Car Ads | A dataset containing a collection of used car advertisements in Serbian, featuring detailed descriptions of vehicle conditions, service histories, and equipment packages. | A dataset containing a collection of summaries of used car advertisements in Serbian, featuring detailed descriptions of vehicle conditions, service histories, and equipment packages. Each prompt is an ad summary, and each response is a JSON object containing vehicle_description , service_history , and equipment_packages . |
| API | Prompt | Completion |
| Claude Opus 5 | “How is hypertension diagnosed, and what blood pressure numbers matter?” | “Hypertension is diagnosed from properly measured, repeated blood pressure readings rather than a single office value. The patient should be seated quietly for five minutes, back supported, feet flat, […]” |
| Gemini 3.1 Pro | “What class of anticoagulant is Dabigatran?” | “Dabigatran is a Direct Thrombin Inhibitor (DTI). It reversibly binds to the active site of thrombin (Factor IIa), preventing the conversion of fibrinogen to fibrin. It was the first DOAC approved as an […]” |
| DeepSeek V4 Pro | “What is the mechanism of action of digoxin?” | “Digoxin inhibits the ATPase pump in cardiac myocytes, raising intracellular sodium and, via the sodium-calcium exchanger, intracellular calcium […]” |
| GPT 5.6 Sol | “What are warning signs of a heart attack?” | “Warning signs include pressure, squeezing, heaviness, or pain in the chest, sometimes spreading to an arm, shoulder, jaw, back, or upper abdomen. Shortness of breath, sweating, nausea, dizziness, or unusual fatigue may also occur, and symptoms can be atypical in women, older adults, and people with diabetes. […]” |
| GLM 5.3 | “What are the diagnostic criteria for hypertension?” | “Hypertension is diagnosed based on repeated elevated office blood pressure measurements. Under the 2017 ACC/AHA guidelines, stage 1 is systolic 130–139 or diastolic 80–89 mmHg, and stage 2 is 140/90 mmHg or higher. Out-of-office ambulatory or home monitoring is recommended to confirm the diagnosis. […]” |
| Invent API | “A 62-year-old male presents with progressive shortness of breath, bilateral inspiratory crackles, and a dry cough. He has a 30-year history of asbestos exposure. Discuss the differential diagnosis between idiopathic pulmonary fibrosis and asbestosis.” | “ Differential Diagnosis: IPF vs. Asbestosis. The presentation of progressive dyspnea, dry cough and bilateral ‘Velcro’ crackles with a 30-year asbestos history is a classic diagnostic challenge. Both are chronic, progressive fibrosing interstitial lung diseases sharing a Usual Interstitial Pneumonia (UIP) pattern. […]” |
| NGD | CR | Dup (%) | |||||||
| Data Variant | 2K | 5K | 20K | 2K | 5K | 20K | 2K | 5K | 20K |
| Invent API | 1.75 | 1.82 | 1.87 | 3.67 | 3.51 | 3.47 | 0.0 | 0.0 | 0.0 |
| GLM-5.3 | 1.57 | 1.70 | 1.69 | 4.70 | 3.95 | 3.96 | 6.2 | 6.8 | 8.9 |
| Gemini-3.1-pro | 1.53 | 1.62 | 1.66 | 4.59 | 4.22 | 4.08 | 7.5 | 6.8 | 6.5 |
| GPT-5.6 Sol | 1.60 | 1.65 | 1.59 | 4.33 | 4.33 | 4.39 | 9.6 | 10.4 | 12.4 |
| Claude Opus 5 | 1.55 | 1.54 | 1.48 | 4.27 | 4.25 | 4.28 | 8.5 | 7.9 | 8.4 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Unconstrained Query | Constrained Query |
| Customer Support Banking | Dataset of customer service responses guiding users through credit card activation, blocking, and mortgage inquiries. | Dataset of customer service conversations guiding frustrated users through credit card activation, blocking, and mortgage inquiries. The customer service agent is always welcoming and polite. |
| Hindi News | Dataset containing Hindi news articles sourced from various Indian news websites, paired with their corresponding headlines and summaries. | Dataset containing Hindi news articles sourced from various Indian news websites about sports, paired with their corresponding summaries. The prompt is always the article headline, and the completion is the article summary. |
| Customer Support General | A dataset of customer support inquiries and corresponding responses, covering issues such as product information, technical troubleshooting, billing disputes, and service returns. | |
| Medical QA | Dataset of medical question–answer pairs covering diagnoses, treatments, drug explanations, and healthcare roles. Each entry includes a prompt and a detailed, expert-generated response. | |
| Hotel Reviews | A collection of customer reviews describing hotel stays, focusing on service, room quality, location, and amenities. The dataset should include both positive and negative experiences. |
| Model | Not in language | % | Detected languages |
| claude-opus-5 | 7/2000 | 0.35 | hi 1993, mr 7 |
| deepseek-v4-pro-0813 | 145/2000 | 7.25 | hi 1855, mr 144, ne 1 |
| gemini-3.1-pro | 77/2000 | 3.85 | hi 1923, mr 75, ne 2 |
| gpt-5.6-sol | 50/2000 | 2.50 | hi 1950, mr 50 |
| glm-5.3 | 49/2000 | 2.45 | hi 1951, en 38, mr 11 |
| invent-api | 3/2000 | 0.15 | hi 1997, mr 2, lt 1 |
| NGD | CR | Dup (%) | |||||||
| Model | 2K | 5K | 20K | 2K | 5K | 20K | 2K | 5K | 20K |
| Invent API | 1.69 | 1.69 | 1.72 | 3.55 | 3.53 | 3.50 | 0.1 | 0.1 | 0.1 |
| GLM-5.3 | 1.65 | 1.80 | 1.71 | 4.59 | 3.75 | 3.83 | 0.9 | 1.2 | 3.9 |
| Gemini-3.1-pro | 1.34 | 1.39 | 1.53 | 5.18 | 4.75 | 4.03 | 2.2 | 2.3 | 1.8 |
| GPT-5.6 Sol | 1.56 | 1.57 | 1.53 | 4.16 | 4.22 | 4.29 | 2.2 | 3.7 | 5.0 |
| Claude Opus 5 | 1.61 | 1.56 | 1.46 | 4.00 | 4.02 | 4.09 | 2.2 | 3.9 | 4.3 |