AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
- Type
- paper
- Venue
- arXiv:2609.38288 (cs.AI), submitted 29 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Targets the self-improving capability of LLM agents — iteratively refining a solution at test time — which rests on two complementary abilities: reflection (produce a solution better than the current one) and long-horizon execution (keep the iteration effective over many rounds). Both are treated as domain-agnostic and learned in scenarios suited to supervision: long-horizon improvement trajectories are synthesized from ML and algorithmic programming tasks, two domains offering verifiable feedback and reward for sustained iteration. The agent, built on Qwen3.8-27B, scores 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, transfers to deep research (84.0 BrowseComp, 52.6 HLE, 92.2 GAIA, 93.8 DeepSearchQA), and keeps improving as its budget of rounds grows — evidence that long-horizon reflective data is an effective route toward self-improving agents.
Keywords
AREX-2 · self-improving agents · long-horizon · reflection · test-time iteration · deep research · MLE-bench · BrowseComp · GAIA
Topics
self-improving agents, long-horizon tasks, reflection, test-time refinement, deep research, Qwen3.8-27B
Research notes
- Discovery: @arXivBangers X post 2026-10-02 2:11 PM ET (https://x.com/arXivBangers/status/2106084605577380147)
- AREX Team, Beijing Academy of Artificial Intelligence (BAAI)
- Code: https://github.com/VectorSpaceLab/AREX-2
- Models: https://huggingface.co/collections/BAAI/arex-2
- Post's link pointed to the arxivb.org mirror of arXiv 2609.38288