We present daVinci-MagiHuman, an open-source audio-video generative foundation model for human-centric generation. daVinci-MagiHuman jointly generates synchronized video and audio using a single-stream Transformer that processes text, video, and audio within a unified token sequence via self-attention only. This single-stream design avoids the complexity of multi-stream or cross-attention architectures while remaining easy to optimize with standard training and inference infrastructure. The model is particularly strong in human-centric scenarios, producing expressive facial performance, natural speech-expression coordination, realistic body motion, and precise audio-video synchronization. It supports multilingual spoken generation across Chinese (Mandarin and Cantonese), English, Japanese, Korean, German, and French. For efficient inference, we combine the single-stream backbone with model distillation, latent-space super-resolution, and a Turbo VAE decoder, enabling generation of a 5-second 256p video in 2 seconds on a single H100 GPU. In automatic evaluation, daVinci-MagiHuman achieves the highest visual quality and text alignment among leading open models, along with the lowest word error rate (14.60%) for speech intelligibility. In pairwise human evaluation, it achieves win rates of 80.0% against Ovi 1.1 and 60.9% against LTX 2.3 over 2000 comparisons. We open-source the complete model stack, including the base model, the distilled model, the super-resolution model, and the inference codebase.
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
We present daVinci-MagiHuman, an open-source audio-video generative foundation model for human-centric generation.
- Year
- 2026
- Venue
- arXiv 2026
- Stars
- 2.0k
- Authors
- 45
- Hosting
- Abstract onlyARXIV-DEFAULT
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2603.21986ARXIV-DEFAULT
- TL;DR
- Semantic Scholar
Abstract
Authors
45Hao WangPengFei LiuWeixian XuTiantian MiYixiu LiuLyumanshan YeYue CaoZheng ZhangSand. aiHansi TengHongyu JiaLingzhi LiTianning ZhangXiaoyang KangYunpeng HuangYutong LinZewei TaoZhongshu WangHanwen SunHong PanMin HuXiaojie CaiLijie LiuSteffi ChernEthan ChernZinan GuoJiadi SuZhulin HuYan MaWenqiang ZhangSII-GAIRJin LiJunjie YuQiangang WangQuanwei QiTao BuTaoran WangTeren XuWentai ZhangXianping YiYunbo ZhangZhaoliang LiuZhiyao CenZhixuan YuZijin Zhou