← Back to explorer

Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs

Type
other
Venue
arXiv / Meta / Harvard

Summary

Diagnoses score dilution in static self-attention: the target–distractor logit margin must scale as Ω(log T) or needle mass vanishes. Thinking tokens cannot recover buried evidence under that bound. qTTT does one prefill to cache K/V then a few gradient steps on W_Q only, raising the margin without invalidating the cache. FLOP-matched to 8k thinking tokens, Qwen3-4B gains +12.6 / +14.1 pp average on LongBench-v2 and ZeroScrolls subsets; thinking-token gains saturate as T grows.

Keywords

qttt · test-time-training · long-context · score-dilution · longbench · zeroscrolls · meta · harvard

Topics

long-context LLMs, test-time training, attention

Research notes

  • Primary: arxiv abs (cs.LG; also cs.CL). CC BY 4.0 on HTML. Work done while at Meta; affiliations Harvard/Kempner, Meta, OpenAI, UC Berkeley, UT Austin. Correspondence rachitbansal@g.harvard.edu, az@astonzhang.com. No official code on abs. HF paper page 2 upvotes; no linked models/datasets. Discord posted abs (Substack UTM). Uses public LongBench-v2/ZeroScrolls; no new corpus, so no datasets_local row. License field left blank per catalog convention.