Adversaries can still steal reasoning from American frontier models via Third-Party Cloud Aggregators ("Stolen Thoughts" update)
- Type
- paper
- Venue
- September 2026 (updated paper; stolen-thoughts.com)
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
Updated paper on reasoning extraction: patching your own API does not secure your cloud-hosting ecosystem. Most critical findings: reasoning replay attacks still intermittently work on Microsoft Azure for every OpenAI model (Astra included) and for Anthropic up to Sonnet 5, with one decode returning the trace verbatim; scratchpad reasoning attacks (ask the model to put its reasoning into a tool argument, publicly demonstrated by @_can1357) still extract reasoning from every OpenAI model and from Opus 4.6 and Sonnet 5; today only Opus 4.1, Fable 3, and Fable 5 do not expose their reasoning. A September 13 audit found replay extraction blocked on direct OpenAI/Anthropic APIs but still working on Azure for the same models -- same models, different defenses, depending on which platform serves them. OpenAI publicly disclosed a reasoning-extraction campaign citing this work: operators copied encrypted reasoning from one conversation and asked a model in another to decrypt and transcribe it; activity began July 1, spiked July 24-25 with 16,000 requests using an extraction pattern from over 4,000 users; related prompt-pattern activity across a cluster of 15,000+ users was fully disrupted by July 28. The authors argue patches must cover all attacks and hosting clouds, or attackers choose the route where protection is weakest; findings shared to accelerate defense rollout.
Keywords
AI safety · reasoning extraction · distillation · model theft
Topics
AI safety, reasoning extraction, distillation, model theft
Research notes
- Discovery: Joachim Schaeffer (@JSchaeff3r, verified, AI safety / AI control researcher) 2026-09-30 thread: https://x.com/JSchaeff3r/status/2105356987630543057?s=20 ("We stole reasoning. Again.")
- Full update PDF: https://stolen-thoughts.com/stolen_thoughts_update.pdf
- OpenAI disclosure: https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/
- No license stated.
- Connects to the collection's AI-safety, model-theft, and distillation entries.