← Back to explorer

Adversaries can still steal reasoning from American frontier models via Third-Party Cloud Aggregators ("Stolen Thoughts" update)

Type
paper
Venue
September 2026 (updated paper; stolen-thoughts.com)
Year
2026
Source
paper
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Updated paper on reasoning extraction: patching your own API does not secure your cloud-hosting ecosystem. Most critical findings: reasoning replay attacks still intermittently work on Microsoft Azure for every OpenAI model (Astra included) and for Anthropic up to Sonnet 5, with one decode returning the trace verbatim; scratchpad reasoning attacks (ask the model to put its reasoning into a tool argument, publicly demonstrated by @_can1357) still extract reasoning from every OpenAI model and from Opus 4.6 and Sonnet 5; today only Opus 4.1, Fable 3, and Fable 5 do not expose their reasoning. A September 13 audit found replay extraction blocked on direct OpenAI/Anthropic APIs but still working on Azure for the same models -- same models, different defenses, depending on which platform serves them. OpenAI publicly disclosed a reasoning-extraction campaign citing this work: operators copied encrypted reasoning from one conversation and asked a model in another to decrypt and transcribe it; activity began July 1, spiked July 24-25 with 16,000 requests using an extraction pattern from over 4,000 users; related prompt-pattern activity across a cluster of 15,000+ users was fully disrupted by July 28. The authors argue patches must cover all attacks and hosting clouds, or attackers choose the route where protection is weakest; findings shared to accelerate defense rollout.

Keywords

AI safety · reasoning extraction · distillation · model theft

Topics

AI safety, reasoning extraction, distillation, model theft

Research notes

  • Discovery: Joachim Schaeffer (@JSchaeff3r, verified, AI safety / AI control researcher) 2026-09-30 thread: https://x.com/JSchaeff3r/status/2105356987630543057?s=20 ("We stole reasoning. Again.")
  • Full update PDF: https://stolen-thoughts.com/stolen_thoughts_update.pdf
  • OpenAI disclosure: https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/
  • No license stated.
  • Connects to the collection's AI-safety, model-theft, and distillation entries.