Post-Training Leaves Behavioral Shadows on Unrelated Decisions
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
Finds that language models can transfer capabilities through task-unrelated text. Active Taskless Distillation (ATD) achieves capability transfer using only a single word from the teacher per prompt: it selects prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words, and a student initialized from that ancestor learns solely from the resulting prompt-word pairs, with no target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664 instances yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control; transfer also appears in scientific knowledge, commonsense reasoning, and reading comprehension across model generations, sizes, and families. The learned shadow is composable and its strength tracks the teacher's update strength.
Keywords
distillation · post-training · subliminal learning · capabilities · interpretability
Topics
distillation, post-training, subliminal learning, interpretability
Research notes
- User-supplied: https://huggingface.co/papers/2609.29233
- Submitted 2026-09-24, cs.CL.
- Connects to the collection's distillation, subliminal-learning, and interpretability entries.