← Back to explorer

RL is an evolutionary algorithm

Type
blog
Venue
snimu.github.io (personal blog)
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Argues that pretraining, RL, and compaction are all evolutionary algorithms: micro-batch SGD is evolutionary search in the loss landscape (only generalizing updates survive), agent compaction evolves lesson summaries under environment feedback, and RL is evolutionary search in the reward landscape (weights as species, behaviors as individuals, reward as selection). Draws implications for generalization (multi-domain RL beats multi-teacher distillation), honesty/monitorability, prompt-reward design (mismatch breeds eval awareness), and agentic judging as a path to align capabilities with alignment.

Keywords

reinforcement learning · evolutionary algorithms · pretraining · alignment · agent design · agentic judges

Topics

reinforcement learning, evolutionary algorithms, pretraining, alignment, agent design

Research notes

  • Discovery: shared in #random-papers on 2026-09-29.
  • Method: Conceptual essay framing pretraining (micro-batch SGD as evolutionary search in the loss landscape), compaction (agent lesson summaries evolving under environment feedback), and RL (weights as species, sampled behaviors as individuals, GRPO groups as populations, reward as selection) as evolutionary algorithms.
  • Key findings: multi-domain RL should generalize better than multi-teacher on-policy distillation; instruction-reward mismatch drives eval awareness as an evolved pattern; agentic judges can make alignment part of the environment.
  • Conclusion: environment and agentic judge design are the two most impactful areas of both capabilities and alignment research.