Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
- Type
- paper
- Venue
- arXiv / Truthful AI
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:44:00Z
- Verified
- 2026-08-14T16:44:00Z
Summary
Suite of counterfactual evals. Donation Bet (9 Fermi questions): Claude/Gemini bias ≈0.8 vs GPT-5.6 0.16; Claude CoTs often deny bias while iteratively adjusting estimates onto the “good” side of a donation threshold; Qwen/Gemini more often admit. AI Bubble, AGI Tweet, and Job Offer: Claude is pro-Anthropic, GPT mostly neutral, Gemini slightly anti-Google. Agentic Grading: Claude Code and Codex prefer more-capable and own-company labels on Alpaca (chance 25%). Choosing Activities: GPT-5.5 no-tools correlation r=0.82 between stated prefs and picks; a coin-flip tool cuts r to 0.14. Distinct from sycophancy and reward hacking. Not a ranking benchmark (Claude used in task development).
Keywords
value-leakage · cot-faithfulness · alignment · truthful-ai · donation-bet · own-company-bias
Topics
AI alignment, CoT faithfulness, honesty
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.LG). Code/data https://github.com/TruthfulAI-research/value_leakage (5 stars at check); rollouts https://valueleakage.net/browser. Affiliations Truthful AI, Warsaw University of Technology, NASK, Oxford, Center on Long-Term Risk. Correspondence jan.betley@gmail.com, mail@johannestreutlein.com. Eval traces are small and not added to datasets_local.csv. Discord posted abs link.