← Back to explorer

Invent a Dataset: Training Data for Custom Models

Type
blog
Venue
Adaption blog
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Adaption (founded by ex-Cohere/Google researchers Sara Hooker and Sudip Roy) launched Invent a Dataset: describe the behavior you want a custom model to learn in plain English and get back a structured, training-ready dataset — no seed corpus, schema design, or labeling required. One datasets.invent API call sets domains, row count, format, and language expansion; generation is async and output is instruction pairs or preference pairs downloadable as JSONL, JSON, CSV, or Parquet. It is the first half of a zero-data loop with AutoScientist (launched May 2026), which co-optimizes the data and training recipe against the objective — Adaption reports AutoScientist beats human-configured training by 35% on average (win rates 48% to 64% across eight verticals, 5k-100k rows, on Together AI fine-tuning architectures).

Keywords

synthetic data · dataset generation · custom models · AutoScientist · Adaption

Topics

synthetic data, dataset generation, custom models, AutoScientist, Adaption

Research notes

  • Method: Intent-driven synthetic data generation: the API interprets a natural-language objective, defines the dataset structure, and generates training examples at requested size/scope; dataset IDs feed directly into autoscientist.create to close the intent-to-trained-model loop.
  • Key findings: In-house evaluations: AutoScientist-configured training beats staff-configured training by 35% on average across eight domain-specialized verticals
  • Limitations: Performance figures are in-house and unaudited; the product is a commercial API (ADAPTION_API_KEY required), not an open artifact.
  • Blog page not fetched directly (rate-limited); substance verified via the official docs (docs.adaptionlabs.ai), the API docs repo, and secondary coverage (runtimewire, marktechpost).