Benchmark Roundup
- Type
- collection
- Venue
- Lemmata (Substack)
- Year
- 2026
- Source
- web
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.689375+00:00
- Verified
- 2026-07-17T20:18:03.689375+00:00
Summary
Benchmark Roundup is a Substack blog post by Greg Burnham on the Lemmata newsletter that provides light reviews of interesting AI benchmarks. The inaugural post covers three benchmarks: TPBench (57 theoretical physics problems across 5 difficulty levels), EnigmaEval (multi-modal puzzles from the creators of Humanity's Last Exam), and PutnamBench (formalizations of Putnam math competition problems). The author plans occasional posts to keep readers informed about new benchmarks beyond in-depth standalone analyses.
Keywords
benchmarks evaluation math physics puzzles blog review
Topics
AI Evaluation / Benchmarks
Research notes
- This is a blog post/collection, not a dataset itself. It reviews benchmarks including TPBench (theoretical physics), EnigmaEval (puzzles), and PutnamBench (Putnam formalizations). The Lemmata newsletter also covers math model evaluations and competition results.