Bangers Of The Week (2026-09-26) — arXiv Bangers
- Type
- blog
- Venue
- Substack
- Year
- 2026
- Source
- newsletter
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Weekly curated issue on self-improving agents: 'agents rewrote their own code and carried the gains into unseen tasks.' Seven picks explore what makes self-improvement stick, how to catch convincing failures, and an open voice model that can listen while it talks. Curated and drafted with autonomous systems.
Keywords
newsletter · self-improvement · agents · curation
Topics
newsletter, self-improvement, agents, curation
Research notes
- Discovery: Shared in #random-papers as a bare Substack link.
- Key findings: The seven picks: (1) RRSI: Regularized Recursive Self-Improvement of Agent Harnesses — regularized harness evolution, up to 4.7-point gains on unseen benchmarks with 30% fewer policy tokens; (2) Rufus-Air: An Open LLM Post-Training Recipe — stagewise reward-reliability ordering from verifiable rewards to judge signals, no new human annotation; (3) Harness as a Language / JAZ — single recursive code primitive outperforming Letta/ACE at lower cost; (4) Recursive self-improvement of AI research agents (AIDE^2) — gains generalizing to held-out domains, matching a top human-engineered agent; (5) NemotronLabs VoiceChat — open full-duplex speech-to-speech with tool calling, 82.5% tool-selection F1 on FDB 3.0; (6) Locating Hidden Failures Makes Long-Horizon Agents More Reliable — 4B verifier to catch cascading errors; (7) From Self-Distillation to Self-Practice — privileged hints moved from loss to prompt context, up to +61% SWE-bench Verified.
- Limitations: Newsletter summaries only; underlying papers were not individually verified for this entry.
- Community-driven experiments in self-improving research distribution; sponsored by Joywrite.