Dolci-Think-SFT-32B
- Type
- dataset
- Venue
- Allen AI (AI2) / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English (primary), with some multilingual content (Aya dataset includes multiple languages)
- Added
- 2026-07-17T20:18:03.700641+00:00
- Verified
- 2026-07-17T20:18:03.700641+00:00
Summary
Dolci-Think-SFT-32B is a 2.25M-row SFT dataset assembled by AI2 for training the Olmo-3-32B-Think reasoning model. It combines existing reasoning traces from OpenThoughts 3 (941K prompts, Apache 2.0), SYNTHETIC-2 (105K prompts), and NVIDIA Nemotron Post-Training code split (114K prompts), with new reasoning traces generated by DeepSeek R1 and DeepSeek R1 0528 for repurposed Tülu 3/OLMo 2 prompts including WildChat (76K), WildJailbreak (40K), Aya (97K), WildGuardMix (37K), CoCoNot (10K), OpenAssistant Guanaco (7K), and TableGPT (5K). The dataset underwent extensive filtering for data quality and topic safety via Azure API, and is part of the Olmo 3 training pipeline (SFT → DPO → RLVR).
Keywords
reasoning sft olmo ai2 chain-of-thought deepseek-r1 tulu fine-tuning
Topics
NLP / Reasoning
Research notes
- The 32B variant is at allenai/Dolci-Think-SFT-32B; a 7B variant exists at allenai/Dolci-Think-SFT-7B with slightly different composition. Associated with arxiv: 2507.02833 and 2512.13961. Sources are a mixture of Apache 2.0, ODC-BY, CC-BY-4.0, and MIT licensed data.