Large AI models are driving surging demand for computing power. Alibaba Cloud’s new Wanxiang 2.6 series not only delivers significant improvements in image quality, sound effects, and instruction-following capabilities but also extends single-video generation length to a domestic-leading 15 seconds—achieved through efficient optimization of computing infrastructure. The model family now integrates over ten visual creation capabilities, including text-to-image, image editing, text-to-video, image-to-video, human-voice-to-video, motion generation, role-playing, and general video editing. This versatile toolchain for AI content generation demonstrates the broad potential of large models in the creative industry. The core technological breakthrough of Wanxiang 2.6 lies in its industry-first “role-playing” function in China. Using multimodal joint modeling, it extracts a character’s appearance, voice, emotion, pose, visual features, and acoustic characteristics from input videos, ensuring consistent transfer across all sensory dimensions. The model generates dynamic content featuring single persons, multiple persons, or person-object interactions. Additionally, a new professional storyboard control feature converts user prompts into multi-scene scripts, constructing coherent narrative videos. High-level semantic understanding guarantees subject and scene consistency during camera transitions, reflecting cutting-edge advances in the convergence of natural language processing and computer vision.
文章图片 2
Empowering professional film and TV creation, Wanxiang 2.6’s role-playing and storyboard control greatly simplify production workflows. Ordinary users can upload personal videos, input prompts in a sci-fi mystery style, and the model completes storyboard design, character performance, and dubbing within minutes, producing short films with complete narrative and cinematic camera movements—realizing the creative dream of “everyone a director.” For professional fields such as advertising design and short drama production, continuous prompt input generates complete narrative shorts, lowering creative barriers and promoting AI adoption in the creative industry. This signals a clear trend toward more efficient and specialized AI applications. Since Alibaba first introduced the synchronized audio-video Wanxiang 2.5 domestically in September, the series has ranked first in China on the authoritative large model evaluation set LMArena. The launch of version 2.6 further solidifies Alibaba’s leading position in video generation large models. Wanxiang 2.6 is now available on Alibaba Cloud’s Bailian platform and the Wanxiang official website for public testing; enterprise users can access it via API. The Qianwen APP also plans to integrate the model soon, offering richer interactive experiences. This highlights the critical role of chip infrastructure in enabling large model deployment, accelerating the transition from AI research to industrial application.