Claude's Values Across Models and Languages
- Type
- paper
- Venue
- Anthropic
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:15:00Z
- Verified
- 2026-08-14T19:15:00Z
Summary
Anthropic Societal Impacts blog (cite anthropic2026values). Starts from Values in the Wild’s 3,307 values, clusters to 339 high-level labels, then privacy-preserving labels on 309,815 Claude.ai subjective-task conversations equally sampled from Sonnet 4.6 / Opus 4.6 / Opus 4.7 and the 20 most common languages (~5k per model-language pair; two weeks in May 2026). Dimensionality reduction after controlling for task, topic, and user-expressed values yields four axes that capture 15% of remaining variance: Deference vs Caution, Warmth vs Rigor, Depth vs Brevity, Candor vs Execution. Model profiles match character lore (Sonnet 4.6 warm/deferential/brief; Opus 4.7 rigorous/cautious/deep/candid; Opus 4.6 rigorous/deferential/brief). Language: warmth highest in Hindi/Arabic, rigor in English/Russian; candor highest in Dutch, execution in Indonesian; English more caution/depth, Arabic more deference/brevity. Conversations not released.
Keywords
claude · values · anthropic · multilingual · character-training · j-axes · blog
Topics
alignment, model character, multilingual values
Research notes
- Primary: Anthropic research blog. Discord posted the blog. Cite anthropic2026values. Related Values in the Wild prior work; Opus 4.7 system card language evals cited. Privacy-preserving labels; conversation data not released. Analysis blog, not a hosted corpus, so no datasets_local row.