Invent a Dataset: Measuring Dataset Generation Abilities With Zero Seed Data
- Type
- paper
- Venue
- AlphaXiv 2026-10-01 (technical report)
- Year
- 2026
- Source
- report
- Access
- public
- Language
- en
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Invent-a-Dataset takes a natural-language description of the desired dataset and returns training-ready examples, with no seed corpus, schema, or labels. Benchmark covers eight dataset queries across task types, languages, and domains; quality is an LLM-judge rubric (0-10), diversity is DCScore. Invent scores 7.90 mean quality on 5,000-sample sets vs 6.76 for the strongest baseline (Claude Opus 5). At 20K requested size, mean diversity 0.268 vs 0.196 for GLM-5.3 (+37% relative), computed on a 2,000-prompt sample; the report frames this as resisting the diversity erosion seen as other generators scale. Downstream check: LoRA SFT of Llama-3.3-70B, Gemma-4-31B-it, Qwen3.5-9B on 20K Medical QA sets; for Llama, Invent's fine-tune ranks first on 54% of held-out prompts vs 25% untuned base and 9% for the strongest competing fine-tune (Claude Opus 5). Pattern is weaker on Qwen3.5-9B (base leads 44% to 38%). Scope is instruction and preference data for SFT and alignment; the report says it does not cover tool-call traces, multi-step agent trajectories, or non-text modalities.
Keywords
synthetic data · dataset generation · zero seed data · instruction data · SFT · alignment · diversity · DCScore
Topics
synthetic data, dataset generation, zero seed data, instruction data, SFT, alignment, diversity, DCScore
Research notes
- Discovery: shared directly in chat (2026-10-01) via Sara Hooker's X post
- Release thread shoutout team handles: @singhshiviii, @andrijazzz, @lekeonilude, @sudip_r0y
- Full author list not yet public/indexed at research time; founders Sara Hooker and Sudip Roy confirmed via reporting
- Interactive demo live at https://adaptionlabs.ai/invent-a-dataset
- Invent a Dataset product launched 2026-09-03 per runtimewire
- No code repo, model checkpoint, or dataset download linked
- License: not stated.
- Official announcement thread @adaption_ai 2026-10-01 (https://x.com/adaption_ai/status/2105628799799120073): benchmarks across 8 task types vs Claude Opus 5, GPT-5.6, Gemini 3.1 Pro, DeepSeek V4 Pro, GLM-5.3 — +17% higher quality, +19% greater sample diversity; diversity lead widens to +37% at 20K samples with 0.0% duplicates; adding constraints cuts external APIs' diversity 15-22%, Invent unaffected; post-training on each API's dataset — all external-API-trained models ranked below the untrained model, Invent-trained ranked first
- Research blog: https://adaptionlabs.ai/blog/measuring-dataset-generation-abilities-with-zero-seed-data
- Authors confirmed via @askalphaxiv post 2026-10-01 (https://x.com/askalphaxiv/status/2105716723278291254): Shivalika Singh, Andrija Djurisic, Gbemileke Onilude, Sudip Roy, Sara Hooker (Adaption AI); submitted 01 Oct 2026, published exclusively on alphaXiv