MemoHarness: Agent Harnesses That Learn from Experience
- Type
- paper
- Venue
- arXiv / University of Notre Dame
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:44:00Z
- Verified
- 2026-08-14T16:44:00Z
Summary
Decomposes the harness into six control surfaces (context, tool, generation, orchestration, memory, output) plus a dual-layer experience bank of per-case diagnoses and distilled global patterns. Training-time search (T=10) starts from a minimal harness with GPT-5.3-Codex; test-time adaptation retrieves similar successes/failures with no labels or extra search. On Terminal-Bench 18-task held-out split, 0.806 vs Codex 0.722 (+0.084); LiveCodeBench 0.900→0.967; FinanceAgent 0.600→0.767. Cross-model mean +0.098 (GLM-5 +0.233). Reported cost $6.89 vs Codex $10.28 because 13.32M of 14.18M input tokens were cached. Authors note the 18-task split, no CIs, and incomplete component ablations.
Keywords
memoharness · agent-harness · terminal-bench · experience-bank · test-time-adaptation · notre-dame
Topics
LLM agents, harness engineering
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.AI). Code https://github.com/HowieHwong/MemoHarness (88 stars at check). Notre Dame / LMU Munich / USC; corresponding xzhang33@nd.edu. HF paper page 2 upvotes. Discord posted abs link. Paper not a dataset; no datasets_local row.