RL is an evolutionary algorithm
- Type
- blog
- Venue
- snimu.github.io (personal blog)
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Argues that pretraining, RL, and compaction are all evolutionary algorithms: micro-batch SGD is evolutionary search in the loss landscape (only generalizing updates survive), agent compaction evolves lesson summaries under environment feedback, and RL is evolutionary search in the reward landscape (weights as species, behaviors as individuals, reward as selection). Draws implications for generalization (multi-domain RL beats multi-teacher distillation), honesty/monitorability, prompt-reward design (mismatch breeds eval awareness), and agentic judging as a path to align capabilities with alignment.
Keywords
reinforcement learning · evolutionary algorithms · pretraining · alignment · agent design · agentic judges
Topics
reinforcement learning, evolutionary algorithms, pretraining, alignment, agent design
Research notes
- Discovery: shared in #random-papers on 2026-09-29.
- Method: Conceptual essay framing pretraining (micro-batch SGD as evolutionary search in the loss landscape), compaction (agent lesson summaries evolving under environment feedback), and RL (weights as species, sampled behaviors as individuals, GRPO groups as populations, reward as selection) as evolutionary algorithms.
- Key findings: multi-domain RL should generalize better than multi-teacher on-policy distillation; instruction-reward mismatch drives eval awareness as an evolved pattern; agentic judges can make alignment part of the environment.
- Conclusion: environment and agentic judge design are the two most impactful areas of both capabilities and alignment research.