Invent a Dataset: Training Data for Custom Models
- Type
- blog
- Venue
- Adaption blog
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Adaption (founded by ex-Cohere/Google researchers Sara Hooker and Sudip Roy) launched Invent a Dataset: describe the behavior you want a custom model to learn in plain English and get back a structured, training-ready dataset — no seed corpus, schema design, or labeling required. One datasets.invent API call sets domains, row count, format, and language expansion; generation is async and output is instruction pairs or preference pairs downloadable as JSONL, JSON, CSV, or Parquet. It is the first half of a zero-data loop with AutoScientist (launched May 2026), which co-optimizes the data and training recipe against the objective — Adaption reports AutoScientist beats human-configured training by 35% on average (win rates 48% to 64% across eight verticals, 5k-100k rows, on Together AI fine-tuning architectures).
Keywords
synthetic data · dataset generation · custom models · AutoScientist · Adaption
Topics
synthetic data, dataset generation, custom models, AutoScientist, Adaption
Research notes
- Method: Intent-driven synthetic data generation: the API interprets a natural-language objective, defines the dataset structure, and generates training examples at requested size/scope; dataset IDs feed directly into autoscientist.create to close the intent-to-trained-model loop.
- Key findings: In-house evaluations: AutoScientist-configured training beats staff-configured training by 35% on average across eight domain-specialized verticals
- Limitations: Performance figures are in-house and unaudited; the product is a commercial API (ADAPTION_API_KEY required), not an open artifact.
- Blog page not fetched directly (rate-limited); substance verified via the official docs (docs.adaptionlabs.ai), the API docs repo, and secondary coverage (runtimewire, marktechpost).