← Back to explorer

Liquid: Language Models are Scalable and Unified Multi-modal Generators

Type
other
Venue
arXiv / Huazhong University of Science and Technology / ByteDance / The University of Hong Kong

Summary

Tokenizes images into discrete codes and trains them with text tokens in a single LLM, with no CLIP encoder or external diffusion head. Reports a scaling law: the usual unified-training drop shrinks as size grows from 0.5B to 32B (Gemma-2 2B/9B, Qwen2.5 0.5B/7B/32B). Continue-training existing LLMs (Gemma-7B for the main Liquid-7B) is claimed 100× cheaper than from-scratch and beats Chameleon while staying near Llama 2 on text. Liquid-7B FID 5.47 on MJHQ-30K, below the listed autoregressive unified models and most diffusion baselines including SD-XL. Code https://github.com/FoundationVision/Liquid; project https://foundationvision.github.io/Liquid/.

Keywords

liquid · unified-mllm · discrete-codes · autoregressive-generation · bytedance · hust · mjhq · chameleon

Topics

multimodal generation, unified models, autoregressive image generation

Research notes

  • Primary: arxiv abs (cs.CV). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: Technical report; project https://foundationvision.github.io/Liquid/. HUST / ByteDance / HKU. Correspondence xbai@hust.edu.cn, songbai.site@gmail.com; Yi Jiang project lead. Code https://github.com/FoundationVision/Liquid (642 stars at check). HF paper page 3 upvotes; githubRepo linked; official-looking linked model Junfeng5/Liquid_V1_7B not copied into hf_* fields. Discord posted abs. Training mix is JourneyDB/synthetic T2I plus public text, not a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.