How to Steal Reasoning Without Reasoning Traces
- Type
- paper
- Venue
- arXiv / Cornell Tech
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T20:50:00Z
- Verified
- 2026-08-14T20:50:00Z
Summary
Cornell Tech Trace Inversion attack: train an inversion model on a weaker surrogate (e.g. R1-Distill-Qwen-1.5B) that sees full traces, then invert a black-box victim's (x, answer, optional reasoning summary) into synthetic long CoT. Inverted traces overlap R1 ground truth (TF1 52.76 with weak surrogate; ~81% length recovery) and teach students better than answer-only, summary+answer, or surrogate traces. Fine-tuning Qwen2.5-7B-Instruct on traces inverted from GPT-5.4 mini summaries lifts JEEBench 19.7% to 31.6% vs surrogate traces; Llama-3.1-8B MATH500 16.4% to 52.4% vs summary+answer. Hiding traces / antidistillation sampling does not block this because inversion ignores the victim's internal reasoning. Code Apache-2.0.
Keywords
trace-inversion · distillation · cot · model-stealing · reasoning · cornell-tech
Topics
LLM security, distillation, chain-of-thought, model stealing
Research notes
- Primary: arxiv abs v2 (cs.CR). Resolved from Discord HF papers URL huggingface.co/papers/2603.07267. Code https://github.com/Tingwei-Zhang/Trace_Inversion_Attack (Apache-2.0). HF paper page also links unofficial Jackrong Trace-Inverter / Negentropy models and datasets — not official artifacts. Later Discord also posted the abs (1536350933519175781); this row uses the original HF-papers message.