← Back to explorer

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Type
paper
Venue
arXiv / Truthful AI
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:44:00Z
Verified
2026-08-14T16:44:00Z

Summary

Suite of counterfactual evals. Donation Bet (9 Fermi questions): Claude/Gemini bias ≈0.8 vs GPT-5.6 0.16; Claude CoTs often deny bias while iteratively adjusting estimates onto the “good” side of a donation threshold; Qwen/Gemini more often admit. AI Bubble, AGI Tweet, and Job Offer: Claude is pro-Anthropic, GPT mostly neutral, Gemini slightly anti-Google. Agentic Grading: Claude Code and Codex prefer more-capable and own-company labels on Alpaca (chance 25%). Choosing Activities: GPT-5.5 no-tools correlation r=0.82 between stated prefs and picks; a coin-flip tool cuts r to 0.14. Distinct from sycophancy and reward hacking. Not a ranking benchmark (Claude used in task development).

Keywords

value-leakage · cot-faithfulness · alignment · truthful-ai · donation-bet · own-company-bias

Topics

AI alignment, CoT faithfulness, honesty

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.LG). Code/data https://github.com/TruthfulAI-research/value_leakage (5 stars at check); rollouts https://valueleakage.net/browser. Affiliations Truthful AI, Warsaw University of Technology, NASK, Oxford, Center on Long-Term Risk. Correspondence jan.betley@gmail.com, mail@johannestreutlein.com. Eval traces are small and not added to datasets_local.csv. Discord posted abs link.