AutoBenchmark: benchmark creation and the role of humans (Meta RAM blog)
- Type
- blog
- Venue
- Meta AI, RAM blog (September 2026)
- Year
- 2026
- Source
- blog
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
AutoBenchmark is a framework in which an autoresearch agent builds a benchmark end-to-end from a task specification and revises it over iterations using two feedback signals: solver-agent trajectories and scores, and critique from an LLM verifier checking benchmark quality. Each benchmark is produced as a Harbor-compatible evaluation package (container environment, task instructions, evidence, reference solution, machine-checkable grading criteria). Stages: 1) Benchmark Proposal -- agent decides the construct, operationalizes it multiple ways, gathers primary sources, instantiates runnable tasks, makes the designer decisions (what the solver sees, partial credit, reference solution proving solvability); 2) Benchmark Solving -- solver agents attempt tasks, graded by an answer judge, giving a difficulty signal; 3) Benchmark Review -- an LLM judge rates the benchmark on construct validity, correctness, feasibility, usefulness, and overall verdict, retaining the best accepted checkpoint (lowest-scoring accepted iteration) and checking difficulty against external solvers. Before any solver run, the harness executes the agent's reference solution against the real verifier and admits it only at score >= 0.9. Three benchmarks built automatically: Graveyard Bench (ideation -- avoiding dead-end ideas, grounded in documented negative results), SilentTrain Bench (experimentation -- patching buggy code that silently degrades training, grounded in silent defects), Rebuttal Bench (assessment -- judging whether a paper rebuttal resolves reviewer weaknesses, grounded in public review threads). Findings: fully autonomous benchmarks are close to saturated (in-loop solvers above 80, first three iterations all 100.0 on Muse Spark); fine-grained human proposal feedback halves solver scores (Rebuttal Bench: 90/84/80 -> 43.5/51.2/39.1 on Spark/Glimmer/Nemotron); coarse one-sentence intent helps only marginally; difficulty transfers to held-out solvers (NVIDIA-Nemotron-3.5-Lightning-30B-A3B, Claude Opus-5); when the loop stalls, human direction at that point helps (SilentTrain: 88.1 -> 64.2 on Spark); the verifier catches defects scores cannot see (leaked answers, shallow constructs), so selection on score alone keeps verifier-rejected iterations.
Keywords
benchmarks · autoresearch · human-in-the-loop · agents · Harbor
Topics
benchmark creation, autoresearch, human-in-the-loop, agents
Research notes
- Discovery: Jason Weston (@jaseweston, verified, Meta Senior Director) 2026-09-30 thread: https://x.com/jaseweston/status/2105305463784935791
- Takeaways: human-agent collaboration gives big wins over agents alone; fine-grained feedback in ideation is crucial; the recipe enables autoresearch benchmarks for AI research (full recursive improvement loop).
- Full technical report coming to arXiv "soon" -- no paper link yet.
- Connects to the collection's autoresearch, benchmark-creation, Harbor, and agent-evaluation entries.