SenseNova-U1-8B-MoT
- Type
- model
- Venue
- SenseTime / SenseNova
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:32:00Z
- Verified
- 2026-08-14T18:32:00Z
Summary
HF weights for SenseNova-U1-8B-MoT: a dense unified multimodal model (NEO-unify / NEOChat) that drops the visual encoder and VAE and uses Mixture-of-Transformers experts for understanding vs generation. Card: ~8B understanding + ~8B generation params (inspect script: 17.552B total; understanding 8.121B / generation 8.186B / shared embed+lm_head 1.245B). Native interleaved image-text, infographic T2I, editing, VQA; siblings include SFT, Infographic, 8-step preview, LoRA, and A3B-MoT. Code https://github.com/OpenSenseNova/SenseNova-U1. Paper arXiv 2605.12500. Apache-2.0.
Keywords
sensenova-u1 · neo-unify · mot · multimodal · text-to-image · interleaved · sensetime · huggingface
Topics
unified multimodal, image generation, vision-language
Research notes
- Primary: HF model card (apache-2.0, transformers, pipeline_tag any-to-any, downloads=28336 likes=291 checked 2026-08-14). Paper arXiv 2605.12500 (cs.CV; Atom API has no license). Official code https://github.com/OpenSenseNova/SenseNova-U1. Collection https://huggingface.co/collections/sensenova/sensenova-u1. Discord posted the HF model pointer. Weights/checkpoint item, not a new hosted pretraining corpus, so no datasets_local row. Unofficial GGUF sibling smthem/SenseNova-U1-8B-MoT-Merger-gguf not copied into hf_* fields.