← Back to explorer

Goedel-Prover-V2 SFT Dataset

Type
dataset
Venue
Goedel-LM
Year
2026
Source
huggingface
Access
free
Language
English, Lean 4
Added
2026-07-17T20:18:03.736910+00:00
Verified
2026-07-17T20:18:03.736910+00:00

Summary

The Goedel-Prover-V2 SFT Dataset contains 1,745,010 samples of Lean 4 theorem proving supervised fine-tuning data, used to train Goedel-Prover-V2, the strongest open-source theorem prover. It includes synthetic proof tasks generated via scaffolded data synthesis (increasing difficulty), with verifier-guided self-correction using Lean compiler feedback. The resulting 32B model achieves 88.1% pass@32 on MiniF2F and 90.4% in self-correction mode.

Keywords

theorem-proving lean4 formal-math synthetic sft automated-reasoning math proof

Topics

Math / Theorem Proving

Research notes

  • Paper: arxiv 2508.03613. Goedel-Prover-V2-32B ranks first among open-source models on PutnamBench (86 problems at pass@184).