← Back to explorer

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

Type
paper
Venue
arXiv / OpenAI
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T19:06:00Z
Verified
2026-08-14T19:06:00Z

Summary

OpenAI trains with 5% synthetic beneficial-trait conversations (honesty, corrigibility, fairness, etc. across 12 domains) mixed into standard RL. Vs compute-matched baseline: held-out trait score 0.406→0.607; 44/53 OOD alignment evals improve (mean +9.1 pp). Health-only 5% still lifts non-health evals; excluding health/science still lifts health evals. More resistant to harmful persona prompts and (vs pre-RL) harmful medical finetuning. Generic-helpfulness rewards on the same chats do not reproduce the gains. Closed data and model.

Keywords

alignment · beneficial-rl · openai · emergent-misalignment · corrigibility · health · blog

Topics

alignment, RLHF/RL, emergent misalignment

Research notes

  • Primary: arxiv abs 2606.24014. Discord posted https://alignment.openai.com/beneficial-rl/. Correspondence ajag@openai.com, karan@openai.com. Equal-contribution Jagadeesh/Singhal. Synthetic trait dataset not released. Not a hosted corpus, so no datasets_local row.