OpenThoughts-Agent: Data Recipes for Agentic Models
- Type
- paper
- Venue
- arXiv (cs.AI)
- Year
- 2026
- Source
- x
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Addresses the lack of public knowledge on data curation for agentic models via 100+ controlled ablations on a six-stage SFT curation pipeline plus agentic RL data studies. Fine-tuning Qwen3-32B on the resulting 100K-example OpenThoughts-Agent-v2 yields 44.8% average across seven agentic benchmarks (+3.9pp over the strongest open-data baseline), with strong data scaling properties.
Research notes
- Key findings: Task source and diversity matter most in the SFT curation pipeline; OpenThinker-Agent-32B is the strongest open-data <=32B model on the seven-benchmark average (44.8%, +3.9pp over best baseline); SFT and RL stages can compose (8B-scale evidence)
- Limitations: RL study only at 8B scale; no base-model ablation (all Qwen3); largest set 100K trajectories - extrapolation to multi-million-trajectory regimes untested.
- Announced by Alex Dimakis on X 2026-06-24; blog published 2026-06-10 at open-thoughts.ai/blog/agent.