Harness Learning Enables Generalizable Test-Time Adaptation
- Type
- paper
- Venue
- arXiv:2609.35738 (cs.CL), submitted 28 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Introduces harness learning: instead of updating model weights at test time, train a proposer model to revise the solver's executable harness — the program organizing model calls, tool use, and information flow — using execution feedback. Framed as meta-learning over executable programs, with harness revisions playing the role of weight updates; the proposer is trained with RL using revised harnesses' task performance as reward, and candidate harnesses run under a frozen solver. On 21 unseen Reasoning Gym families, SFT followed by RL raises mean single-revision score from 0.32 (base proposer) to 0.62, and the trained 4B proposer beats its 35B teacher on average. A QA proposer trained with RL only on HotpotQA transfers to MuSiQue and 2WikiMultihopQA, with ten revision rounds nearly doubling exact match. Policies trained on single revisions keep improving harnesses over multiple rounds, suggesting harness design can become a reusable adaptation skill.
Keywords
harness learning · harness optimization · test-time adaptation · meta-learning · proposer model · execution feedback · Reasoning Gym · HotpotQA · MuSiQue · continual learning agents
Topics
agent harnesses, test-time adaptation, meta-learning, reinforcement learning, multi-hop QA, continual learning
Research notes
- Discovery: @rsalakhu X post 2026-10-01 (https://x.com/rsalakhu/status/2105730919466246640) quoting lead author @ZAlvin39105's thread 2026-09-30 (https://x.com/ZAlvin39105/status/2105329046359875658)
- Outer/inner loop: given task description, current harness, and execution report, the proposer writes a code edit; candidate harnesses run with a frozen solver and their task performance is the RL reward; at test time the inner loop repeatedly revises the harness from execution feedback while both models' weights stay fixed
- Reasoning: 21 unseen Reasoning Gym families (excluded from both training phases); SFT + RL raises mean single-revision score 0.32 -> 0.62; trained 4B proposer outperforms its 35B teacher on average (0.62 vs 0.56), though the teacher remains stronger under oracle best-of-eight selection; gains are in the mean over all proposals, so individual revisions become more useful before best-candidate selection
- QA: separate proposer trained directly with RL on HotpotQA (no teacher demonstrations) transfers to MuSiQue and 2WikiMultihopQA; on MuSiQue ten revision rounds raise exact match ~0.14 -> 0.27 (avg of four runs); on both unseen QA benchmarks each RL proposer's average final run exceeds the oracle best of 80 independent revisions
- What revisions learn: e.g. keep summaries for guiding retrieval but pass original retrieved passages to the answering call; add retrieval hops or re-search when a draft answer is absent from passages
- Open questions noted: training on revision sequences does not consistently improve over single-revision training (per-edit immediate rewards leave cross-chain credit assignment unexplored); SFT data from a single teacher over a limited range of harness designs; fixed context-assembly procedure limits failure patterns seen
- Code (release soon): https://github.com/AlvinZH04/Harness-Learning
- Website: https://alvinzh04.github.io/Harness-Learning-Website/
- License: CC BY 4.0