← Back to explorer

Inverse RL Helps Align AI by Imitating Humans

Type
paper
Venue
arXiv / Carnegie Mellon University / UT Austin
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:40:00Z
Verified
2026-08-14T16:40:00Z

Summary

PARED (Projected Alignment Reward Estimated from Demonstrations) trains a lightweight logistic discriminator that separates expert demonstrations from policy samples in a practitioner-chosen response-level feature space (Gemma-3-27B helpfulness/harmlessness scores plus five LDA topics; length excluded as a shortcut). No task-specific preference labels. As inference-time best-of-16 on GPT-OSS-20B, wins 63.4% of non-tied Gemini-2.5-Flash judgments. On-policy GRPO of Qwen2.5-3B-Instruct: with 4,000 audience-conditioned HH demos, ab-initio PARED reaches 84.6% vs Instruct (SFT 81.1%) and post-hoc PARED wins 86–88% vs its SFT init; with 500 demos, Instruct-init PARED 70.7% vs SFT 58.0%. Separate adult/child rewards improve both audiences. Demonstrations are prompted GPT-5.1 completions, not human data. No official code on the abs page.

Keywords

pared · inverse-rl · alignment · demonstrations · grpo · contextual-alignment · hh-rlhf · qwen2.5

Topics

LLM alignment, inverse RL, contextual alignment

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). Affiliations CMU, UT Austin, independent. Correspondence mwilinsk@cs.cmu.edu. No official code on abs. Discord posted abs link.