← Back to explorer

ReToken: One Token to Improve Vision–Language Models for Visual Retrieval

Type
paper
Venue
arXiv / UIUC / Microsoft Research
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:32:00Z
Verified
2026-08-14T16:32:00Z

Summary

Diagnoses query–key attention as a weak visual retriever in VLMs and scores frames in value space instead. ReToken is one learnable token plus a final-layer projection, trained with class-balanced BCE on MIRAGE multi-image QA while the VLM is frozen by default. On Visual Haystacks it lifts Qwen3VL-8B by 13.4 points and InternVL3.5-8B by 12.4 points at C=50 (>20% relative); trained only on images it transfers zero-shot to hour-scale video (+8.0 on LVBench with Qwen3VL-8B). Training and long-video inference fit on one H100. Code https://github.com/avaxiao/ReToken.

Keywords

retoken · visual-retrieval · kv-cache · vlm · qwen3vl · internvl · visual-haystacks · lvbench · value-space

Topics

vision-language models, visual retrieval, long-context VLMs

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.CV/AI/LG). Code MIT https://github.com/avaxiao/ReToken (13 stars at check). Affiliations UIUC, Microsoft Research, Google DeepMind (Zhu; work done at UIUC). Correspondence yaox11@illinois.edu, dhoiem@illinois.edu. Discord posted AlphaXiv 2607.28627.