Reinforcement Learning Towards Broadly and Persistently Beneficial Models
- Type
- paper
- Venue
- arXiv / OpenAI
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:06:00Z
- Verified
- 2026-08-14T19:06:00Z
Summary
OpenAI trains with 5% synthetic beneficial-trait conversations (honesty, corrigibility, fairness, etc. across 12 domains) mixed into standard RL. Vs compute-matched baseline: held-out trait score 0.406→0.607; 44/53 OOD alignment evals improve (mean +9.1 pp). Health-only 5% still lifts non-health evals; excluding health/science still lifts health evals. More resistant to harmful persona prompts and (vs pre-RL) harmful medical finetuning. Generic-helpfulness rewards on the same chats do not reproduce the gains. Closed data and model.
Keywords
alignment · beneficial-rl · openai · emergent-misalignment · corrigibility · health · blog
Topics
alignment, RLHF/RL, emergent misalignment
Research notes
- Primary: arxiv abs 2606.24014. Discord posted https://alignment.openai.com/beneficial-rl/. Correspondence ajag@openai.com, karan@openai.com. Equal-contribution Jagadeesh/Singhal. Synthetic trait dataset not released. Not a hosted corpus, so no datasets_local row.