← Back to explorer

Bangers Of The Week (2026-09-26) — arXiv Bangers

Type
blog
Venue
Substack
Year
2026
Source
newsletter
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Weekly curated issue on self-improving agents: 'agents rewrote their own code and carried the gains into unseen tasks.' Seven picks explore what makes self-improvement stick, how to catch convincing failures, and an open voice model that can listen while it talks. Curated and drafted with autonomous systems.

Keywords

newsletter · self-improvement · agents · curation

Topics

newsletter, self-improvement, agents, curation

Research notes

  • Discovery: Shared in #random-papers as a bare Substack link.
  • Key findings: The seven picks: (1) RRSI: Regularized Recursive Self-Improvement of Agent Harnesses — regularized harness evolution, up to 4.7-point gains on unseen benchmarks with 30% fewer policy tokens; (2) Rufus-Air: An Open LLM Post-Training Recipe — stagewise reward-reliability ordering from verifiable rewards to judge signals, no new human annotation; (3) Harness as a Language / JAZ — single recursive code primitive outperforming Letta/ACE at lower cost; (4) Recursive self-improvement of AI research agents (AIDE^2) — gains generalizing to held-out domains, matching a top human-engineered agent; (5) NemotronLabs VoiceChat — open full-duplex speech-to-speech with tool calling, 82.5% tool-selection F1 on FDB 3.0; (6) Locating Hidden Failures Makes Long-Horizon Agents More Reliable — 4B verifier to catch cascading errors; (7) From Self-Distillation to Self-Practice — privileged hints moved from loss to prompt context, up to +61% SWE-bench Verified.
  • Limitations: Newsletter summaries only; underlying papers were not individually verified for this entry.
  • Community-driven experiments in self-improving research distribution; sponsored by Joywrite.