← Back to explorer

Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect

Type
paper
Venue
arXiv / Harvard College
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:44:00Z
Verified
2026-08-14T16:44:00Z

Summary

Shows binary yes/no injection detection is confounded in small models (r=0.999 with a factual-no control on Llama-3.1-8B). Replaces it with sentence localization (N=10, chance 10%) and strength comparison (chance 50%). Across Llama-3.2 1B/3B/8B and Gemma-4 2B/4B/26B-A4B, introspection emerges at 2–3B and generally rises with scale; Llama-1B is at or below chance. IFT (LoRA r=16 on the model’s own steered forward passes) lifts Llama-1B average localization 9.6%→60.6% (semantic, random layer) and strength comparison 30.2%→52.2% zero-shot; Llama-3B loc 14.4%→34.7%; Llama-8B 16.8%→28.3%. MMLU/Winogrande mostly preserved (Llama-1B MMLU 49.1→46.7 under Random·Semantic).

Keywords

ift · introspection · activation-steering · llama · gemma · harvard

Topics

interpretability, introspection, activation steering

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.CL). Code https://anonymous.4open.science/r/IFT-introspection-2092/README.md (anonymous at catalog time). Harvard College. HF has no paper page (404). Discord posted a tweet plus PDF. Concept sets are generated with GPT-4.1 Mini and are not a standalone public corpus; no datasets_local row.