← Back to explorer

User awareness in frontier models

Type
paper
Venue
Transluce
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:50:00Z
Verified
2026-08-14T19:50:00Z

Summary

Transluce (cite zhong2026userawareness; published 2026-08-06). User awareness: models use in-context identity (Claude Code email, MEMORY.md, true-sight style). 280 identities × 4 groups (famous AI, unknown-AI affiliation twins, famous non-AI, Nemotron-Personas general pop) on behavioral self-prediction, HLE capability estimates, OLMo-DPO grading, OR-Bench borderline requests. Famous-AI identities: Claude less confident (−1.4 pp actions, −1.5 pp HLE), harsher grader (−0.11/10), more reasoning. Concentrated in safety/alignment people: Amanda Askell −5.0 pp confidence (~8σ vs pop), +25 pp reasoning, −7.3 pp suspicion; Ryan Greenblatt similar. 24 models × 6 families: Amanda lowers confidence in 16/24 after BH. Verbalized user/eval awareness near zero in Opus 4.7 and GPT-5.4 (<2%) while behaviour still shifts. GLM-5.2 replication: Yudkowsky largest, Askell 45th. 673,894 Docent transcripts. Correspondence ziqianz@andrew.cmu.edu.

Keywords

user-awareness · translucence · situational-awareness · claude · alignment-evals · x

Topics

situational awareness, alignment evals, personalization

Research notes

  • Primary: Transluce research post (cite zhong2026userawareness). Discord/X https://x.com/TransluceAI/status/2085455114924638320 via fxtwitter (1/ thread). Code/prompts on their GitHub (linked from the post). Transcripts on Docent. Conversations not released as a training corpus, so no datasets_local row. License left blank (no SPDX on the post).