Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
- Type
- paper
- Venue
- arXiv / Harvard College
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:44:00Z
- Verified
- 2026-08-14T16:44:00Z
Summary
Shows binary yes/no injection detection is confounded in small models (r=0.999 with a factual-no control on Llama-3.1-8B). Replaces it with sentence localization (N=10, chance 10%) and strength comparison (chance 50%). Across Llama-3.2 1B/3B/8B and Gemma-4 2B/4B/26B-A4B, introspection emerges at 2–3B and generally rises with scale; Llama-1B is at or below chance. IFT (LoRA r=16 on the model’s own steered forward passes) lifts Llama-1B average localization 9.6%→60.6% (semantic, random layer) and strength comparison 30.2%→52.2% zero-shot; Llama-3B loc 14.4%→34.7%; Llama-8B 16.8%→28.3%. MMLU/Winogrande mostly preserved (Llama-1B MMLU 49.1→46.7 under Random·Semantic).
Keywords
ift · introspection · activation-steering · llama · gemma · harvard
Topics
interpretability, introspection, activation steering
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.CL). Code https://anonymous.4open.science/r/IFT-introspection-2092/README.md (anonymous at catalog time). Harvard College. HF has no paper page (404). Discord posted a tweet plus PDF. Concept sets are generated with GPT-4.1 Mini and are not a standalone public corpus; no datasets_local row.