← Back to explorer

Frontier LLMs Still Struggle with Simple Reasoning Tasks

Type
paper
Venue
arXiv / Google DeepMind
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T20:26:00Z
Verified
2026-08-14T20:26:00Z

Summary

DeepMind/Princeton (cite Malek et al.; arXiv 2507.07313). Procedurally generated tasks with tunable tediousness (word/char counting, FOL eval/negation, MathGAP-style proof trees, travel planning) keep the same difficulty while increasing computation; even o1/o3/Gemini 2.5 Pro/R1 degrade as parameters grow (error accumulation, long context, statistical shortcuts, poor state tracking, OOD vocab, copy/tokenization). Unpuzzles: 97 famous puzzles plus minimally edited trivial unpuzzles; models score far higher on hard originals than on easy unpuzzles (gaps 9–54 pp) via memorized solutions and "reasoning delirium." Context-shifted unpuzzles (64 numerical items) restore accuracy, isolating wording/memorization. Complements GSM-Symbolic-style perturbations by decreasing difficulty rather than increasing it.

Keywords

unpuzzles · reasoning-evals · memorization · test-time · ood · deepmind · mathgap

Topics

LLM reasoning evals, memorization, OOD generalization

Research notes

  • Primary: arxiv abs 2507.07313 (cs.CL). Code/data https://github.com/google-deepmind/unpuzzles_and_simple_reasoning (Apache-2.0). Discord posted PDF; a nearby message links the GitHub. Small eval set, not a training corpus. HF paper page missing at fetch.