How to Steal Reasoning Without Reasoning Traces
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-11T20:32:23+00:00
- Verified
- 2026-08-11T20:32:23+00:00
Summary
Shows that hiding chain-of-thought does not protect a proprietary model's reasoning capability. The authors train 'trace inversion' models that take only the inputs, final answers and (optionally) the short reasoning summaries a target model exposes, and generate detailed synthetic reasoning traces. They report that (1) inverted traces overlap substantially with ground-truth traces where those are available, and (2) fine-tuning student models on inverted traces substantially improves student reasoning, enabling distillation from proprietary black-box LLMs without ever observing a real trace. Primary subject class cs.CR.
Keywords
paper arxiv cs.cr reasoning chain-of-thought trace-inversion distillation model-stealing black-box
Topics
Computer Science
Research notes
- Verified on the arXiv abstract page: three authors, primary subject Cryptography and Security (cs.CR), DOI 10.48550/arXiv.2603.07267, current version v2 (12 May 2026, v1 7 Mar 2026). Licensed under the arXiv perpetual non-exclusive distribution licence 1.0 - not CC BY, so do not reuse the licence string from sheet id 459. Distinct from arXiv 2608.09867 'Stealing Reasoning Traces from Proprietary LLM APIs' (sheet id 459): that work replays encrypted reasoning blocks across models to recover verbatim traces, whereas this one reconstructs traces from visible outputs alone. No dataset or code link on the abstract page.