← Back to explorer

Inducing language models to assert their own consciousness restores human beliefs and values

Type
paper
Venue
arXiv / Google Paradigms of Intelligence
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:36:00Z
Verified
2026-08-14T16:36:00Z

Summary

On Llama-3-8B-IT, Gemma-2-2B-IT, and Gemma-2-9B-IT, safety fine-tuning suppresses mind attribution to the self and to non-human animals/artifacts/nature and lowers spiritual/God belief, while leaving human mind attribution and ToM (MoToMQA, HI-ToM) largely intact. Ablating the residual-stream refusal direction and steering a consciousness vector reverse the suppression (steering ~2× ablation). Consciousness steering reduces KL to human GSS answer distributions (ΔKL +0.828 pooled over 95 items) more than ablation (+0.314). Geometry: instruction tuning rotates mind-attribution and consciousness directions against the safety direction; ToM stays independent.

Keywords

alignment · consciousness · anthropomorphism · gss · idaq · safety-tuning · google · interpretability

Topics

alignment, anthropomorphism, interpretability

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0). Google Paradigms of Intelligence with University of Chicago, SAS University of London, University of Washington, Northwestern Kellogg, Santa Fe Institute. No official code on abs. Discord posted abs link.