← Back to explorer

How to Steal Reasoning Without Reasoning Traces

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-11T20:32:23+00:00
Verified
2026-08-11T20:32:23+00:00

Summary

Shows that hiding chain-of-thought does not protect a proprietary model's reasoning capability. The authors train 'trace inversion' models that take only the inputs, final answers and (optionally) the short reasoning summaries a target model exposes, and generate detailed synthetic reasoning traces. They report that (1) inverted traces overlap substantially with ground-truth traces where those are available, and (2) fine-tuning student models on inverted traces substantially improves student reasoning, enabling distillation from proprietary black-box LLMs without ever observing a real trace. Primary subject class cs.CR.

Keywords

paper arxiv cs.cr reasoning chain-of-thought trace-inversion distillation model-stealing black-box

Topics

Computer Science

Research notes

  • Verified on the arXiv abstract page: three authors, primary subject Cryptography and Security (cs.CR), DOI 10.48550/arXiv.2603.07267, current version v2 (12 May 2026, v1 7 Mar 2026). Licensed under the arXiv perpetual non-exclusive distribution licence 1.0 - not CC BY, so do not reuse the licence string from sheet id 459. Distinct from arXiv 2608.09867 'Stealing Reasoning Traces from Proprietary LLM APIs' (sheet id 459): that work replays encrypted reasoning blocks across models to recover verbatim traces, whereas this one reconstructs traces from visible outputs alone. No dataset or code link on the abstract page.