Unpuzzles and Simple Reasoning
- Type
- dataset
- Venue
- Google DeepMind
- Year
- 2026
- Source
- github
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.578337+00:00
- Verified
- 2026-07-17T20:18:03.578337+00:00
Summary
This repository provides evaluation data for the paper 'Frontier LLMs Still Struggle with Simple Reasoning Tasks' (Malek et al., 2025) from Google DeepMind. It contains three datasets: (1) 1,120 procedurally generated simple reasoning questions across 5 categories (counting, first-order logic, proof trees, travel planning, etc.) with tunable difficulty parameters; (2) 97 pairs of well-known logic puzzles and their 'unpuzzled' (trivialized) versions; and (3) 69 context-shifted unpuzzles where language/setting is changed but logic preserved. The datasets demonstrate that even frontier thinking models fail on easy problems, exhibiting memorization of original puzzles and systematic failure patterns.
Keywords
reasoning llm-evaluation puzzles benchmark deepmind memorization out-of-distribution synthetic
Topics
Math / Reasoning / LLM Evaluation
Research notes
- GitHub repo is accessible. The paper is available on arXiv (2507.07313). The simple reasoning tasks are procedurally generated with a specific seed, while unpuzzles are manually constructed by minimal textual edits to original puzzles.