← Back to explorer

LibraryDesignBench: can agents design libraries other agents can use?

Type
paper
Venue
arXiv (2026-09)
Year
2026
Source
paper
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Agents will soon replace humans as the main users and designers of libraries; a good library lets future agents write correct programs with less code. LibraryDesignBench: one agent designs a library, and three other agents write programs with it. 15 design tasks across four languages, with no prescribed APIs, abstractions, or guidance, to elicit genuine design decisions; programs are scored on correctness and simplicity. Results: agents using libraries designed by Opus 5.5 or Fable 5.1 pass as many tests as with the production human-written library while writing simpler code; with DeepSeek V4 Pro's libraries, agents do worse than with no library at all. Despite full design freedom, agents converged on the same patterns as the human library for 11 of 15 tasks. Most excess code stems from the library's design, not the implementer: rigid interfaces led agents to reimplement functionality (adding bugs), verbose interfaces forced more code. Giving the designer explicit guidance (start from prototypes, test with fresh subagents, more examples) made downstream programs shorter and improved score modestly at roughly twice the design cost. A standalone spin-off benchmark, LibraryUseBench, tests how well a model uses a real library with minimal guidance: Opus 5.5 leads, Sonnet 5.5 close behind, every newer model improves on its predecessor. Supported by DARPA, NSF, PrimeIntellect, and SnorkelAI's Open Benchmark Grant.

Keywords

benchmarks · agents · library design · API design

Topics

benchmarks, agents, library design, API design

Research notes

  • Discovery: Gabe Orlanski (@GOrlanski, verified, PhD student @WisconsinCS) 2026-09-30 thread: https://x.com/GOrlanski/status/2105353831584235804?s=20
  • Leaderboard: https://ldbench.com
  • Code: https://github.com/SprocketLab/librarydesignbench
  • HF paper page: https://huggingface.co/papers/2609.36730
  • Author blog: https://gabeorlanski.github.io/posts/library-design-bench/
  • Notable thread question (Patch): do libraries designed by one model transfer to agents using other models.
  • License terms not stated in the thread; check paper/repo.
  • Connects to the collection's agent-benchmark and API-design entries.