← Back to explorer

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Type
paper
Venue
arXiv
Year
2026
Source
x
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

The most capable LLaVA-OneVision vision-language model to date. Built on a native OneVision-Encoder with Windowed Attention; codec-stream tokenization treats compressed video as a continuous bit-cost stream where bit-cost dynamics set adaptive temporal groups and motion-residual cues select salient spatial evidence; a shared 3D RoPE unifies codec canvases, sampled frames, and images. Training stack: ~8M re-captioned video samples for pretraining, 4M-sample spatial corpus for fine-tuning. Introduces JumpScore, a temporal-localization benchmark for fine-grained grounding in high-frequency, densely repeated motion.

Keywords

vision-language · video · LLaVA · tokenization · benchmark

Topics

vision-language, video, LLaVA, tokenization, benchmark

Research notes

  • Discovery: Shared in #random-papers as an X post (Brian Bo Li) together with two arXiv links: 2602.08683 (OneVision-Encoder) and 2605.25979 (LLaVA-OneVision-2).
  • Method: Codec-stream tokenization: bit-cost dynamics determine adaptive temporal groups; motion-residual cues select spatial evidence into compact visual canvases; shared 3D RoPE for unified spatiotemporal coordinates.
  • Key findings: LLaVA-OneVision-2-8B reaches 74.9 JumpScore mAP vs Qwen3-VL-8B 30.1 (+44.8). At matched visual-token budgets, codec-stream inputs improve temporal grounding over frame sampling by +9.7 points. Across standard benchmarks, +4.3 average points on video tasks, +5.3 on spatial tasks, +15.6 average J&F on tracking vs Qwen3-VL-8B.
  • Companion paper shared in the same Discord message: OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence (arXiv:2602.08683, Feilong Tang et al., submitted 2026-02-09, revised 2026-02-26). Its Codec Patchification focuses compute on the 3.1%-25% of regions rich in signal entropy; reports +4.1% average over Qwen3-ViT on video understanding tasks across 16 benchmarks with fewer tokens and less pretraining data.