LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
- Type
- paper
- Venue
- arXiv (2026-09-29), cs.CL; CC BY 4.0
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
A benchmark evaluating both the effectiveness and efficiency of long-context LM harnesses (RLMs, ReAct, mini-swe-agent-style coding agents). Existing long-context evaluations are saturated across harnesses, so the four hard tasks require strategic/efficient reasoning (multiple valid strategies with very different costs), hard retrieval (both GREP and semantic), and extensive multi-step reasoning. Example: Constraint Solving asks for every person satisfying 5 conditions scattered across documents; checking the most selective condition first narrows candidates far more cheaply than checking every condition. 5 frontier LMs x 4 agentic harnesses: best setup reaches only 68% macro-average accuracy and no single harness leads across all tasks. Clear accuracy gaps, plus >10x efficiency differences even with the same LM and similar accuracy: more compute does not always mean better accuracy. mini-swe-agent is especially strong, reaching the best performance at a cost close to direct inference. Establishes efficiency as an axis for long-context evaluation and provides a testbed for harnesses that process context strategically rather than exhaustively.
Keywords
long-context · harnesses · agents · benchmarks · efficiency · retrieval
Topics
long-context, harnesses, agents, benchmarks, efficiency
Research notes
- Discovery: Xi Ye (@xiye_nlp, verified, UAlberta/Princeton PLI) 2026-09-30 thread: https://x.com/xiye_nlp/status/2105347196854087985?s=20 (6 posts)
- Project website: https://stringnlplab.github.io/longharness/
- No code repo linked from the post or the arXiv page (Code, Data, Media tab shows only arXivLabs tool toggles).
- License: CC BY 4.0 (via arXiv view-license link).
- Connects to the collection's long-context, agent, and benchmark entries (RLMs, ReAct, mini-swe-agent).