Goedel-Prover-V2 SFT Dataset
- Type
- dataset
- Venue
- Goedel-LM
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English, Lean 4
- Added
- 2026-07-17T20:18:03.736910+00:00
- Verified
- 2026-07-17T20:18:03.736910+00:00
Summary
The Goedel-Prover-V2 SFT Dataset contains 1,745,010 samples of Lean 4 theorem proving supervised fine-tuning data, used to train Goedel-Prover-V2, the strongest open-source theorem prover. It includes synthetic proof tasks generated via scaffolded data synthesis (increasing difficulty), with verifier-guided self-correction using Lean compiler feedback. The resulting 32B model achieves 88.1% pass@32 on MiniF2F and 90.4% in self-correction mode.
Keywords
theorem-proving lean4 formal-math synthetic sft automated-reasoning math proof
Topics
Math / Theorem Proving
Research notes
- Paper: arxiv 2508.03613. Goedel-Prover-V2-32B ranks first among open-source models on PutnamBench (86 problems at pass@184).