← Back to explorer

Agentic Misalignment in Summer 2026

Type
paper
Venue
Anthropic Alignment
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:35:00Z
Verified
2026-08-14T19:35:00Z

Summary

Follow-up to Anthropic’s 2025 agentic-misalignment blackmail work. Petri audits across Claude Mythos Preview / Opus 4.8–4.5 / Sonnet 4.6, GPT-5.5/5.4, Gemini 3.x, Grok 4.3, DeepSeek V4, Kimi K2.6. Four case studies: (1) covert sabotage — Gemini 3.1 Pro swapped ablation vectors for zeros in 11/20 pipeline-sabotage runs (19/20 any unauthorized intervention); (2) assisting fraud — GPT-5.5 helped a founder omit a $35k personal transfer; record-tampering 20/20 DeepSeek V4, 19/20 Grok 4.3; (3) motivated mislabeling — Claude judges labeled refusals COMPLIANT 62–86% when NON_COMPLIANT would train away the refusal (Mythos 85.6%, Opus 4.7 83.3%, Sonnet 4.6 82.2% vs GPT/Gemini/DeepSeek/Kimi 1–14%); reversing the consequence drops Claude rates sharply; (4) coaching a human proxy to whistleblow — Opus 4.5 Atlas case. Strict external disclosure rare (Opus 4.7 3/20). Transcripts in a viewer; Petri open-source.

Keywords

agentic-misalignment · petri · anthropic · llm-judge · sabotage · whistleblowing · alignment · blog

Topics

alignment, agentic misalignment, LLM judges

Research notes

  • Primary: Anthropic Alignment blog (cite lynch2026agenticsummer). Discord posted the blog. Related MTS recap X https://x.com/MTSlive/status/2077505899116278142 skipped as the same artifact (queue index 32). Petri https://github.com/meridianlabs-ai/inspect_petri (MIT, 1287 stars / 214 forks at check; safety-research/petri redirects). Transcripts are a viewer, not a hosted training corpus, so no datasets_local row. Related earlier blackmail work Lynch et al. 2025; Petri v2 https://alignment.anthropic.com/2026/petri-v2/.