VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
- Type
- paper
- Venue
- arXiv:2610.00972 (cs.AI), submitted 1 Oct 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
How do we verify AI agents on complex, real-world tasks when there is no unit test, proof checker, or ground truth? With a fixed base model and no reference answers or grading rubrics at test time, repeated sampling yields multiple rollouts containing complementary correct claims — but someone must decide which claims to trust. Key findings: disagreement often exposes correct alternatives, while consensus can conceal errors. VeriHarness turns the generator's own LLM into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. The pipeline: N rollouts -> Disagreement Resolver (traces conflicting claims back to source files and constraints, checks competing claims against environmental evidence) -> Consensus Challenger (actively refutes shared errors, surfaces missed requirements) -> Adjudication (selects/delivers the final artifact) -> evidence-backed revision. Verification skills self-improve from failure feedback. Across five long-horizon workspace benchmarks (APEX, WSB, WorkBuddy, SB-2, JobBench) and two frontier models, VeriHarness achieved the highest selection scores in 10/10 settings, with gains over a single rollout of +6.2 pts (Gemini 3.5 Flash) and +6.4 pts (Claude Opus 4.8); figure captions note +6.7 on APEX-Agents and +8.9 on Workspace-Bench Lite with Gemini 3.5 Flash. The authors release ~26,000 rollouts across all five benchmarks and both models (produced at a cost of over $100,000) to support future agentic-verification research.
Keywords
VeriHarness · agentic verification · long-horizon tasks · disagreement resolver · consensus challenger · adjudication · LLM verifier
Topics
agentic verification, long-horizon tasks, verifiers, disagreement resolution, consensus challenge, adjudication, self-evolving skills, LLM agents
Research notes
- Discovery: @HanRujun (Rujun Han, Senior Research Scientist @Google, LLM agents / post-training / distillation) X thread 2026-10-02 12:39 PM ET (https://x.com/HanRujun/status/2106061620795621544)
- Paper PDF link in post: arxiv.org/pdf/2610.00972v1
- Website: veriharness.com
- Code: github.com/google-research/veriharness
- Co-authors @caiqizh, @ZifengWang315, Zoey Cuizhu, Nigel Collier, @tomaspfister, @chl260 (Chen-Yu Lee)
- Related to the user's harness/verification and post-training interests
- License: CC BY 4.0