← Back to explorer

Dolci-Think-SFT-32B

Type
dataset
Venue
Allen AI (AI2) / Hugging Face
Year
2026
Source
huggingface
Access
free
Language
English (primary), with some multilingual content (Aya dataset includes multiple languages)
Added
2026-07-17T20:18:03.700641+00:00
Verified
2026-07-17T20:18:03.700641+00:00

Summary

Dolci-Think-SFT-32B is a 2.25M-row SFT dataset assembled by AI2 for training the Olmo-3-32B-Think reasoning model. It combines existing reasoning traces from OpenThoughts 3 (941K prompts, Apache 2.0), SYNTHETIC-2 (105K prompts), and NVIDIA Nemotron Post-Training code split (114K prompts), with new reasoning traces generated by DeepSeek R1 and DeepSeek R1 0528 for repurposed Tülu 3/OLMo 2 prompts including WildChat (76K), WildJailbreak (40K), Aya (97K), WildGuardMix (37K), CoCoNot (10K), OpenAssistant Guanaco (7K), and TableGPT (5K). The dataset underwent extensive filtering for data quality and topic safety via Azure API, and is part of the Olmo 3 training pipeline (SFT → DPO → RLVR).

Keywords

reasoning sft olmo ai2 chain-of-thought deepseek-r1 tulu fine-tuning

Topics

NLP / Reasoning

Research notes

  • The 32B variant is at allenai/Dolci-Think-SFT-32B; a 7B variant exists at allenai/Dolci-Think-SFT-7B with slightly different composition. Associated with arxiv: 2507.02833 and 2512.13961. Sources are a mixture of Apache 2.0, ODC-BY, CC-BY-4.0, and MIT licensed data.