← Back to explorer

Introducing APEX-Agents 1.1

Type
benchmark
Venue
Mercor blog
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

APEX-Agents 1.1 is a revised benchmark testing models' ability to do real professional work across investment banking, management consulting, and corporate law, with task requirements scattered across files and communications. The update targets 'scattergunning' (models hedging with multiple answers to game binary rubric grading) via expert task audits, a scattergun-aware judge model, and explicit anti-hedging prompts. Claude Fable 5.1 leads at 68.6% Pass@1.

Keywords

benchmark · agents · professional work · evaluation · reward hacking

Topics

benchmark, agents, professional work, evaluation, reward hacking

Research notes

  • Method: 80 tasks per domain refined through three expert audits; judge model is DeepSeek-v4-Flash-0731 (temp 0.1) with GEPA-optimized prompt, trained on 1,407 hand-labeled rubric items; scattergunned rubric items scored as zero.
  • Key findings: Claude Fable 5.1 leads at 68.6% Pass@1; GPT-6 Astra most consistent with highest Pass^4 (56.3%); Scattergunning penalties cut some older models' Pass@1 by over 20%; Models that searched communications under ambiguity saw large uplifts (e.g., Kimi K3 +29.7% task-level Pass@1 when it searched)
  • Limitations: Judge optimization raised false-negative rate from 5.3% to 8.0% vs naive grader; results are Mercor's own evaluation.
  • Dataset on Hugging Face and Harbor Hub; eval agent on GitHub (links in post). Blog post date not shown; on/before 2026-09-08.