ProgramBench: write-a-program-from-scratch SWE-agent benchmark; Muse Spark 1.3 results
- Type
- benchmark
- Venue
- programbench.com (results announced 2026-09-30)
- Year
- 2026
- Source
- x
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
ProgramBench: a SWE-agent benchmark where an agent must write a whole program (sqlite, ffmpeg, php, cmatrix, ngrrram, ...) from scratch. Meta Muse Spark 1.3 ranks #2 (max tier) and #3 (xhigh tier) overall on the leaderboard. Full solves by 1.3 max: 5 programs entirely -- cmatrix, eureka, ngrrram*, nomino*, wrapcheck* (* = first time any LM has solved them). ngrrram (a TUI typing trainer): the agent typed lessons into the original and captured the screen frame by frame, then recreated layouts/scoring/UI; remaining misses are small details. Progress 1.1 -> 1.3 (both xhigh): avg test pass rate 47.0% -> 68.6%; almost resolved (>=95%): 8 -> 33 programs; fully resolved: 0 -> 2; 39% fewer tokens across the benchmark. Cost: Muse Spark max averages $6.46/task -- roughly the same as GPT-5.6 Sol xhigh at $6.08, but fully resolves 5 programs vs 2 and gets >=95% on 50 vs 31; Opus 5 reaches 9 solved at $50.53/task (~8x more). Trajectories, final codebases, and test pass/fail breakdowns open-sourced; submissions open to anyone. More models coming soon (6 astra, 5.5 opus noted).
Keywords
agents · coding agents · SWE benchmarks · program synthesis · open-weight evals
Topics
agents, coding agents, SWE benchmarks, program synthesis
Research notes
- Discovery: John Yang (@jyangballin, verified; SWE-bench/SWE-agent author) 2026-09-30 thread: https://x.com/jyangballin/status/2105322290179428389?s=20
- Run page: https://programbench.com/run/muse-spark-1-3-max/
- Trace example: https://programbench.com/trace/muse-spark-1-3-max/wintermute-cell__ngrrram.8ea13c3/
- Submission guide: https://programbench.com/blog/submission-guide/
- Generated codebases: https://huggingface.co/datasets/programbench/20260918_mini-v2.4.5_muse-spark-1-3-max (also cataloged in Datasets).
- No license stated for the open-sourced data.
- Connects to the collection's agent-benchmark/SWE-bench entries.