← Back to explorer

SenseNova-U1-8B-MoT

Type
model
Venue
SenseTime / SenseNova
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-08-14T18:32:00Z
Verified
2026-08-14T18:32:00Z

Summary

HF weights for SenseNova-U1-8B-MoT: a dense unified multimodal model (NEO-unify / NEOChat) that drops the visual encoder and VAE and uses Mixture-of-Transformers experts for understanding vs generation. Card: ~8B understanding + ~8B generation params (inspect script: 17.552B total; understanding 8.121B / generation 8.186B / shared embed+lm_head 1.245B). Native interleaved image-text, infographic T2I, editing, VQA; siblings include SFT, Infographic, 8-step preview, LoRA, and A3B-MoT. Code https://github.com/OpenSenseNova/SenseNova-U1. Paper arXiv 2605.12500. Apache-2.0.

Keywords

sensenova-u1 · neo-unify · mot · multimodal · text-to-image · interleaved · sensetime · huggingface

Topics

unified multimodal, image generation, vision-language

Research notes

  • Primary: HF model card (apache-2.0, transformers, pipeline_tag any-to-any, downloads=28336 likes=291 checked 2026-08-14). Paper arXiv 2605.12500 (cs.CV; Atom API has no license). Official code https://github.com/OpenSenseNova/SenseNova-U1. Collection https://huggingface.co/collections/sensenova/sensenova-u1. Discord posted the HF model pointer. Weights/checkpoint item, not a new hosted pretraining corpus, so no datasets_local row. Unofficial GGUF sibling smthem/SenseNova-U1-8B-MoT-Merger-gguf not copied into hf_* fields.