← Back to explorer

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Finds that language models can transfer capabilities through task-unrelated text. Active Taskless Distillation (ATD) achieves capability transfer using only a single word from the teacher per prompt: it selects prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words, and a student initialized from that ancestor learns solely from the resulting prompt-word pairs, with no target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664 instances yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control; transfer also appears in scientific knowledge, commonsense reasoning, and reading comprehension across model generations, sizes, and families. The learned shadow is composable and its strength tracks the teacher's update strength.

Keywords

distillation · post-training · subliminal learning · capabilities · interpretability

Topics

distillation, post-training, subliminal learning, interpretability

Research notes

  • User-supplied: https://huggingface.co/papers/2609.29233
  • Submitted 2026-09-24, cs.CL.
  • Connects to the collection's distillation, subliminal-learning, and interpretability entries.