Seedance 2.0 is impressive. But it's closed-source!
Introducing our daVinci-MagiHuman — a single-stream 15B Transformer trained from scratch that jointly generates video + audio. No cross-attention. No multi-stream branches. Just self-attention.
5s 1080p video in 38s on a single H100
80% win rate vs Ovi 1.1 | 60.9% vs LTX 2.3 (2,000 human comparisons)
6 languages
Fully open-source
Speed by simplicity.
By @SII_GAIR ×
arxiv.org/abs/2603.21986
github.com/GAIR-NLP/daVin
huggingface.co/spaces/SII-GAI
1:54