HH-RLHF (Human Preference Data for Helpful and Harmless Assistant)
- Type
- dataset
- Venue
- Anthropic
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.804498+00:00
- Verified
- 2026-07-17T20:18:03.804498+00:00
Summary
HH-RLHF is Anthropic's human preference dataset for training helpful and harmless AI assistants via reinforcement learning from human feedback. It contains ~169k chosen/rejected response pairs for helpfulness (from base models, rejection sampling, and online iteration) and harmlessness, plus human-generated red teaming data. The dataset is described in the seminal paper 'Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback' (arXiv:2204.05862).
Keywords
rlhf alignment preference-data helpful harmless red-teaming anthropic reward-model
Topics
NLP / AI Alignment / RLHF
Research notes
- One of the most widely used RLHF preference datasets. GitHub repo (anthropics/hh-rlhf) is deprecated in favor of this HuggingFace dataset. Also includes red teaming data from arXiv:2209.07858. Not intended for supervised training of dialogue agents.