← Back to explorer

OpenThoughts-Agent: Data Recipes for Agentic Models

Type
paper
Venue
arXiv (cs.AI)
Year
2026
Source
x
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Addresses the lack of public knowledge on data curation for agentic models via 100+ controlled ablations on a six-stage SFT curation pipeline plus agentic RL data studies. Fine-tuning Qwen3-32B on the resulting 100K-example OpenThoughts-Agent-v2 yields 44.8% average across seven agentic benchmarks (+3.9pp over the strongest open-data baseline), with strong data scaling properties.

Research notes

  • Key findings: Task source and diversity matter most in the SFT curation pipeline; OpenThinker-Agent-32B is the strongest open-data <=32B model on the seven-benchmark average (44.8%, +3.9pp over best baseline); SFT and RL stages can compose (8B-scale evidence)
  • Limitations: RL study only at 8B scale; no base-model ablation (all Qwen3); largest set 100K trajectories - extrapolation to multi-million-trajectory regimes untested.
  • Announced by Alex Dimakis on X 2026-06-24; blog published 2026-06-10 at open-thoughts.ai/blog/agent.