← Back to explorer

Stealing Reasoning Traces from Proprietary LLM APIs

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-11T20:04:55+00:00
Verified
2026-08-11T20:04:55+00:00

Summary

Leading LLM providers hide chain-of-thought reasoning by returning it to clients as encrypted blocks that are passed back with each request. The authors identify an architectural vulnerability: these encrypted blocks are interchangeable across sessions, users, and models within a provider's ecosystem. Injecting an encrypted reasoning trace from a strong model into a weaker, less safeguarded model of the same provider makes it decode and emit the trace verbatim in plaintext, a scalable "decryption jailbreak" that never jailbreaks the stronger model directly. Four attack vectors are demonstrated across Anthropic, OpenAI, and Google: circumventing anti-distillation protections; large-scale private data extraction (decoding 315,320 reasoning blocks scraped from public repositories recovered 367 PII artifacts and 182 credentials); exposure of hazardous reasoning content even when the visible answer refuses; and invisible prompt injection hidden inside encrypted blocks to poison public agentic rollouts. Cryptographic and system-level mitigations are proposed following responsible disclosure.

Keywords

paper arxiv cs.cr cs.ai cs.lg

Topics

Computer Science