← Back to explorer

Unslopping AI: Reinforcement Learning from eXpert-Aligned Rubrics (RL-XAR)

Type
social
Venue
Meta FAIR (RAM blog)
Year
2026
Source
x
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Introduces RL-XAR (Reinforcement Learning from eXpert-Aligned Rubrics) to fight AI 'slop': learn rubrics that rank expert human writing above model output, then perform iterative RL against those rubrics with periodic rubric re-optimization. Trained Qwen3.5-27B with learned rubrics; Kimi-K2.6 used for rubric meta-optimization. Scientific-section experiments used 561 CS papers / 2,243 training examples and 90 papers / 360 validation examples.

Keywords

reinforcement-learning · alignment · rubrics · writing-quality · meta · reward-hacking

Topics

reinforcement-learning, alignment, rubrics, writing-quality, meta

Research notes

  • Discovery: Posted in #random-papers on 2026-09-28 as an X thread by @swarnaNLP pointing to the FAIR RAM blog post: https://facebookresearch.github.io/RAM/blogs/unslop/ ('Unslopping AI'). Combined into one catalog item as instructed; blog fetched on 2026-09-29.
  • Method: RL-XAR: (1) rubric learning — optimize rubrics so they score expert human writing above model output; (2) iterative RL against learned rubrics; (3) rubric re-optimization loop to stay ahead of reward hacking.
  • Key findings: Reported worst-rubric score of 9.60 with human writing normalized to 10, and a blind expert preference of 16-2 over the baseline model.
  • Limitations: Limitations acknowledged in the post itself: learned rubrics may be biased toward the trained model, and expert-writing measurement is imperfect. Exact authors and publication date not verified; no code or dataset links confirmed.
  • Blog post, not a peer-reviewed paper; treat reported numbers as lab claims. The X thread by @swarnaNLP is the discovery artifact; the blog is the canonical source.