← Back to explorer

Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Type
paper
Venue
arXiv 2609.34378, 2026-09-28
Year
2026
Source
paper
Access
public
Language
en
Added
2026-10-01
Verified
2026-10-01

Summary

Marathoner: Ultra-Long-Horizon Autonomous Intelligence (arXiv 2609.34378; CC BY 4.0). Marathoner-9B, built on Qwen3.5-9B, trained for ultra-long-horizon autonomous coding with a three-stage post-training pipeline: (1) Ultra-Long-Horizon Task Synthesis -- mining major-release PRs (1000+ new lines of code each) from 10,000 GitHub repositories into task-level data in the Harbor framework format, plus 'Multi-Task Chaining' that chains multiple synthesized tasks into one harder task; (2) Rejection Sampling Finetuning -- using Kimi K3 as teacher with diverse harnesses (Claude Code, Codex, etc.), then SFT on rejection-sampled trajectories; (3) Reinforcement Learning -- GRPO in the AReaL framework with real execution in independent sandboxes, plus a novel 'Later Stage Bonus Reward' (+0.5) that explicitly rewards exceptionally valuable operations late in a run (e.g., finding a hidden bug, delivering a major optimization). Endurance: average on FrontierSWE 3.56h / 426.2 steps / 648.3 tool calls; on Terminal-Bench 2.0 0.58h / 104.8 steps / 239.4 tool calls; a case study shows an 11.8-hour trajectory (872 steps) implementing native Rust control-flow operations in Qiskit. Results with Claude Code harness: SWE-bench Verified 43.8% (Qwen3.5-9B) -> 58.4% (after RFT) -> 77.5% (final Marathoner-9B); Terminal-Bench 2.0 27.3% -> 49.7% -> 57.2%; FrontierSWE 10.2 -> 18.2 -> 26.4; NL2Repo 17.9 -> 23.8 -> 34.7; SWE-Marathon 0 -> 2.9 -> 8.2. Compares against proprietary models (Kimi K3, Claude Fable 5.1, GPT-6-Astra, GLM-5.2). Listed under cs.CV; 30 pages, 4 figures. No code, model, or dataset release visible in the paper or arXiv 'Code, Data, Media' section at research time.

Keywords

long-horizon agents · autonomous coding · post-training · rejection sampling · RL · GRPO · Qwen3.5 · SWE-bench Verified · Terminal-Bench 2.0 · FrontierSWE · NL2Repo · tool use

Topics

long-horizon agents, autonomous coding, post-training, rejection sampling, RL, GRPO, Qwen3.5, SWE-bench Verified, Terminal-Bench 2.0, FrontierSWE, tool use

Research notes

  • Discovery: shared directly in chat (2026-10-01) via @mark_k's X post
  • Affiliations: Ant Group, Peking University, University of Macau (FIC)
  • License CC BY 4.0 (arXiv)
  • Tweet cites '100,000 GitHub PRs'; paper specifies 10,000 source repos with major-release PRs mined (1000+ new lines each)
  • PDFs: arxiv.org/pdf/2609.34378, HTML: arxiv.org/html/2609.34378v1, DOI: doi.org/10.48550/arXiv.2609.34378
  • Highly relevant to user's agent work: long-horizon persistence training, multi-task chaining for difficulty, and the later-stage bonus reward technique.