Alibaba has introduced the beta version of Wan3.0, its latest video generation model designed to deliver longer, more consistent, and more versatile AI-generated videos. The new model supports video clips of up to 30 seconds and accepts a broad range of multimodal reference inputs, bringing video creation, visual consistency, and content transformation into a single AI system.

With Wan3.0 now entering public beta, users can apply to test the model through Alibaba Cloud’s AI development platform Model Studio and its AI-native cloud platform Qwen Cloud.

Longer Videos, More Complex Stories

Most mainstream AI video generators produce clips ranging from a few seconds to around 15 seconds. Alibaba’s previous Wan2.7-Video model supported clips of up to 15 seconds, while Wan3.0 doubles that native duration to 30 seconds.

The longer format gives creators more room to execute complex camera movements, continuous shots, and uninterrupted sequences. Wan3.0 also introduces an intelligent duration feature that recommends an appropriate video length based on a user’s prompt.

For creators looking to extend a story further, the model also includes video extension capabilities designed to expand narrative timelines without having to start a new sequence from scratch.

One Model, Multiple Types of Input

One of Wan3.0’s key features is its expanded multimodal input capability. The model can process text, images, videos, and audio simultaneously, while also supporting web pages and documents such as PDFs and PowerPoint presentations.

This allows users to transform traditionally static or text-heavy information into dynamic video content. For businesses, this could make it easier to turn existing documents, presentations, and visual assets into marketing, educational, or explanatory videos.

More Consistent Characters, Products, and Visuals

Wan3.0 is also designed to address a common challenge in AI-generated video: visual drift and distortion as a sequence progresses.

The model focuses on high-precision visual continuity, including realistic human faces, synchronized micro-expressions, multilingual voice outputs, and stable software interfaces and motion graphics.

It can also replicate fine details from reference materials, giving users greater control over characters, products, props, audio, spatial relationships, and visual styles.

Rather than producing approximate visual similarities, Wan3.0 is designed to preserve key elements from reference inputs while maintaining consistent layouts and audio throughout a sequence. These capabilities are combined with natural movement and emotional expressions to create more immersive and cohesive video narratives.

Designed for Creators and Businesses

Wan3.0 is positioned as a versatile tool for a wide range of applications.

Filmmakers and creative teams can use the model to streamline video production, while creators can explore new formats for short dramas and social media content. Businesses can also turn text and images into marketing and educational videos with less manual production work.

The model also has potential applications beyond traditional content creation. Alibaba says Wan3.0 can generate realistic simulation videos that could support technology developers working on areas such as autonomous driving and robotics.

The Evolution of Alibaba’s Wan Series

Alibaba first introduced its Wan series of visual generation models in July 2023. Since then, the company has continued upgrading the technology to make AI-generated images and videos increasingly realistic, controllable, and accessible to creators.

With Wan3.0, Alibaba is pushing the technology further by combining longer video generation, multimodal inputs, stronger visual consistency, and more precise reference control in a single model.

As the model enters public beta, Wan3.0 signals Alibaba’s continued push toward AI-powered creative workflows that can handle not only individual visual assets, but also longer and more complex stories from a wider range of source materials.