Inverse RL Helps Align AI by Imitating Humans
- Type
- paper
- Venue
- arXiv / Carnegie Mellon University / UT Austin
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:40:00Z
- Verified
- 2026-08-14T16:40:00Z
Summary
PARED (Projected Alignment Reward Estimated from Demonstrations) trains a lightweight logistic discriminator that separates expert demonstrations from policy samples in a practitioner-chosen response-level feature space (Gemma-3-27B helpfulness/harmlessness scores plus five LDA topics; length excluded as a shortcut). No task-specific preference labels. As inference-time best-of-16 on GPT-OSS-20B, wins 63.4% of non-tied Gemini-2.5-Flash judgments. On-policy GRPO of Qwen2.5-3B-Instruct: with 4,000 audience-conditioned HH demos, ab-initio PARED reaches 84.6% vs Instruct (SFT 81.1%) and post-hoc PARED wins 86–88% vs its SFT init; with 500 demos, Instruct-init PARED 70.7% vs SFT 58.0%. Separate adult/child rewards improve both audiences. Demonstrations are prompted GPT-5.1 completions, not human data. No official code on the abs page.
Keywords
pared · inverse-rl · alignment · demonstrations · grpo · contextual-alignment · hh-rlhf · qwen2.5
Topics
LLM alignment, inverse RL, contextual alignment
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). Affiliations CMU, UT Austin, independent. Correspondence mwilinsk@cs.cmu.edu. No official code on abs. Discord posted abs link.