Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Type
- other
- Venue
- arXiv / Meta / Harvard
Summary
Diagnoses score dilution in static self-attention: the target–distractor logit margin must scale as Ω(log T) or needle mass vanishes. Thinking tokens cannot recover buried evidence under that bound. qTTT does one prefill to cache K/V then a few gradient steps on W_Q only, raising the margin without invalidating the cache. FLOP-matched to 8k thinking tokens, Qwen3-4B gains +12.6 / +14.1 pp average on LongBench-v2 and ZeroScrolls subsets; thinking-token gains saturate as T grows.
Keywords
qttt · test-time-training · long-context · score-dilution · longbench · zeroscrolls · meta · harvard
Topics
long-context LLMs, test-time training, attention
Research notes
- Primary: arxiv abs (cs.LG; also cs.CL). CC BY 4.0 on HTML. Work done while at Meta; affiliations Harvard/Kempner, Meta, OpenAI, UC Berkeley, UT Austin. Correspondence rachitbansal@g.harvard.edu, az@astonzhang.com. No official code on abs. HF paper page 2 upvotes; no linked models/datasets. Discord posted abs (Substack UTM). Uses public LongBench-v2/ZeroScrolls; no new corpus, so no datasets_local row. License field left blank per catalog convention.