BigO(Bench): Can LLMs Generate Code with Controlled Time and Space Complexity?
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Introduces BigO(Bench), a benchmark for evaluating LLMs on understanding and generating code with specified time/space complexity. Includes tooling to infer algorithmic complexity of any Python function from profiling, plus 3,105 coding problems and 1,190,250 solutions from Code Contests annotated with inferred complexity labels, runtime, and memory footprints across input sizes. Finds token-space reasoning models are unrivaled at code generation but not at complexity understanding, hinting they may not generalize to tasks with no training-time reward.
Keywords
benchmark · code complexity · code generation · Big-O · Code Contests
Topics
code generation, complexity analysis, benchmarking
Research notes
- Discovery: shared by Pierre Chambon (@PierreChambon6, FAIR/Meta AI & INRIA) in an X thread on 2026-09-29 presenting 3 papers on code optimization: https://x.com/PierreChambon6/status/2104966972043657560
- Method: complexity inference tooling from profiling measurements; evaluation of SOTA models on complexity-constrained generation.
- Submitted 2025-03-19 (v1), revised 2026-09-22 (v3).