Inducing language models to assert their own consciousness restores human beliefs and values
- Type
- paper
- Venue
- arXiv / Google Paradigms of Intelligence
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:36:00Z
- Verified
- 2026-08-14T16:36:00Z
Summary
On Llama-3-8B-IT, Gemma-2-2B-IT, and Gemma-2-9B-IT, safety fine-tuning suppresses mind attribution to the self and to non-human animals/artifacts/nature and lowers spiritual/God belief, while leaving human mind attribution and ToM (MoToMQA, HI-ToM) largely intact. Ablating the residual-stream refusal direction and steering a consciousness vector reverse the suppression (steering ~2× ablation). Consciousness steering reduces KL to human GSS answer distributions (ΔKL +0.828 pooled over 95 items) more than ablation (+0.314). Geometry: instruction tuning rotates mind-attribution and consciousness directions against the safety direction; ToM stays independent.
Keywords
alignment · consciousness · anthropomorphism · gss · idaq · safety-tuning · google · interpretability
Topics
alignment, anthropomorphism, interpretability
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0). Google Paradigms of Intelligence with University of Chicago, SAS University of London, University of Washington, Northwestern Kellogg, Santa Fe Institute. No official code on abs. Discord posted abs link.