PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
- Type
- paper
- Venue
- arXiv:2609.38597 (cs.CV), submitted 29 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Presents PixelUMM, an encoder-free unified multimodal model that handles image and video understanding and generation directly in pixel space — no VAE and no ViT. Images are represented as spatial patches and videos as spatiotemporal tubelets, connected to a shared multimodal backbone through single-layer linear projections. Built from Qwen3-8B into a Mixture-of-Transformers, separate understanding and generation experts share one self-attention over text, clean and noisy pixels, jointly supporting autoregressive text prediction and pixel-space flow matching, with clean-pixel prediction extended to video generation. Reported competitive with open-source baselines across image and video understanding and generation, with ablations on decoder design and spatial-temporal patch size. Code and model released.
Keywords
PixelUMM · encoder-free · unified multimodal model · pixel space · Mixture-of-Transformers · tubelets · flow matching · Qwen3-8B · NVIDIA
Topics
unified multimodal models, encoder-free, pixel-space modeling, video generation, Mixture-of-Transformers, flow matching
Research notes
- Discovery: @CongWei1230 X thread 2026-10-01 (https://x.com/CongWei1230/status/2105808995910816159)
- Project page: https://nv-tlabs.github.io/PixelUMM (code and model released via project page, per announcement)
- Motivation: unified multimodal models often use separate visual representations for understanding and generation, lengthening visual context and complicating integration with VLM pretraining; pixel-space modeling is encoder-free but extending it from images to video is non-trivial because video understanding and generation use different temporal representations
- Design (from announcement thread): starts from Qwen3-8B turned into a Mixture-of-Transformers; images -> 16x16 patches, videos -> 16x16x4 tubes, each via one linear layer; separate understanding and generation experts share one self-attention over text, clean and noisy pixels
- Abstract: single-layer linear projections connect raw pixels to the shared backbone; extends clean-pixel prediction to video generation; jointly supports autoregressive text prediction and pixel-space flow matching; competitive performance across image and video understanding and generation; empirical studies of decoder design and spatial-temporal patch size
- Team: NVIDIA NV-T Labs (Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taixé, Sanja Fidler, Wenhu Chen, Zian Wang, Jay Zhangjie Wu)
- License: not stated on arXiv abstract page