← Back to explorer

Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

Type
other
Venue
arXiv / Center for AI Safety

Summary

Fits Thurstonian utilities to forced-choice preferences over 500 world-state outcomes. Preference completeness/transitivity and utility-model accuracy rise with scale (cycle rate <1% at the largest models); linear probes recover utilities in hidden states. Analysis finds political clustering, unequal life-exchange rates, hyperbolic discounting, and decreasing corrigibility. Citizen-assembly SFT on Llama-3.1-8B-Instruct lifts assembly-preference test accuracy 73.2%→90.6% and reduces political bias. Site https://www.emergent-values.ai.

Keywords

utility-engineering · emergent-values · thurstone · citizen-assembly · alignment · cais · hendrycks

Topics

AI values, utility functions, alignment

Research notes

  • Primary: arxiv abs (cs.LG; also cs.AI, cs.CL, cs.CV, cs.CY). License not stated on abs/HTML at check. Center for AI Safety / UPenn / UC Berkeley. Comment: website https://www.emergent-values.ai. No official code on abs. HF paper page 0 upvotes; no linked models/datasets. Discord posted PDF. Preference outcomes and assembly simulation are experimental artifacts, not a hosted corpus, so no datasets_local row. License field left blank per catalog convention.