← Back to explorer

How to Steal Reasoning Without Reasoning Traces

Type
paper
Venue
arXiv / Cornell Tech
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T20:50:00Z
Verified
2026-08-14T20:50:00Z

Summary

Cornell Tech Trace Inversion attack: train an inversion model on a weaker surrogate (e.g. R1-Distill-Qwen-1.5B) that sees full traces, then invert a black-box victim's (x, answer, optional reasoning summary) into synthetic long CoT. Inverted traces overlap R1 ground truth (TF1 52.76 with weak surrogate; ~81% length recovery) and teach students better than answer-only, summary+answer, or surrogate traces. Fine-tuning Qwen2.5-7B-Instruct on traces inverted from GPT-5.4 mini summaries lifts JEEBench 19.7% to 31.6% vs surrogate traces; Llama-3.1-8B MATH500 16.4% to 52.4% vs summary+answer. Hiding traces / antidistillation sampling does not block this because inversion ignores the victim's internal reasoning. Code Apache-2.0.

Keywords

trace-inversion · distillation · cot · model-stealing · reasoning · cornell-tech

Topics

LLM security, distillation, chain-of-thought, model stealing

Research notes

  • Primary: arxiv abs v2 (cs.CR). Resolved from Discord HF papers URL huggingface.co/papers/2603.07267. Code https://github.com/Tingwei-Zhang/Trace_Inversion_Attack (Apache-2.0). HF paper page also links unofficial Jackrong Trace-Inverter / Negentropy models and datasets — not official artifacts. Later Discord also posted the abs (1536350933519175781); this row uses the original HF-papers message.