Hands-On Modern RL
- Type
- repo
- Venue
- WalkingLabs (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Added
- 2026-08-14T18:45:00Z
- Verified
- 2026-08-14T18:45:00Z
Summary
walkinglabs/hands-on-modern-rl is VitePress courseware plus chapter labs: seven parts / 26 chapters from CartPole, MDPs, DQN, policy gradients, PPO, and offline RL into RLHF, DPO, GRPO, RLVR, reasoning models, tool-use/coding/browser/GUI agents, and VLM/audio/embodied RL, with safety/evaluation close. Equations sit next to compact PyTorch. CC BY-NC-SA 4.0. README flags AI-assisted drafting not yet fully reviewed. English translation and PDF via CI (2026-05-15).
Keywords
hands-on-modern-rl · courseware · ppo · dpo · grpo · rlvr · agentic-rl · walkinglabs
Topics
reinforcement learning, LLM alignment, courseware
Research notes
- Primary: GitHub README + LICENSE (CC-BY-NC-SA-4.0; API SPDX NOASSERTION / Other; 3955 stars / 285 forks at check). Course site https://walkinglabs.github.io/hands-on-modern-rl/. Discord posted the repo. README notes AI assistance and incomplete review. Courseware/labs, not a new hosted corpus, so no datasets_local row.