Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Yuan Zhang, Yijun Yang, Hang Xu et al. · 5 May 2026
JoyAI-Image integrates a spatially enhanced MLLM with MMDiT to achieve unified visual understanding, text-to-image generation, and instruction-guided image editing with enhanced spatial intelligence.