← Back to explorer

BigO(Bench): Can LLMs Generate Code with Controlled Time and Space Complexity?

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Introduces BigO(Bench), a benchmark for evaluating LLMs on understanding and generating code with specified time/space complexity. Includes tooling to infer algorithmic complexity of any Python function from profiling, plus 3,105 coding problems and 1,190,250 solutions from Code Contests annotated with inferred complexity labels, runtime, and memory footprints across input sizes. Finds token-space reasoning models are unrivaled at code generation but not at complexity understanding, hinting they may not generalize to tasks with no training-time reward.

Keywords

benchmark · code complexity · code generation · Big-O · Code Contests

Topics

code generation, complexity analysis, benchmarking

Research notes

  • Discovery: shared by Pierre Chambon (@PierreChambon6, FAIR/Meta AI & INRIA) in an X thread on 2026-09-29 presenting 3 papers on code optimization: https://x.com/PierreChambon6/status/2104966972043657560
  • Method: complexity inference tooling from profiling measurements; evaluation of SOTA models on complexity-constrained generation.
  • Submitted 2025-03-19 (v1), revised 2026-09-22 (v3).