← Back to explorer

HH-RLHF (Human Preference Data for Helpful and Harmless Assistant)

Type
dataset
Venue
Anthropic
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.804498+00:00
Verified
2026-07-17T20:18:03.804498+00:00

Summary

HH-RLHF is Anthropic's human preference dataset for training helpful and harmless AI assistants via reinforcement learning from human feedback. It contains ~169k chosen/rejected response pairs for helpfulness (from base models, rejection sampling, and online iteration) and harmlessness, plus human-generated red teaming data. The dataset is described in the seminal paper 'Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback' (arXiv:2204.05862).

Keywords

rlhf alignment preference-data helpful harmless red-teaming anthropic reward-model

Topics

NLP / AI Alignment / RLHF

Research notes

  • One of the most widely used RLHF preference datasets. GitHub repo (anthropics/hh-rlhf) is deprecated in favor of this HuggingFace dataset. Also includes red teaming data from arXiv:2209.07858. Not intended for supervised training of dialogue agents.