Alibaba has launched Wan3.0, the latest generation of its Wan AI video-generation model, with the ability to create videos of up to 30 seconds in a single generation. The model is now available through Alibaba Cloud Model Studio and expands the types of material that can be converted into video, supporting text, images, audio, video and, for the first time in the Wan family, documents such as PDFs, PowerPoint presentations, spreadsheets and other files. Alibaba is positioning the model as a tool for professional creative work, advertising, marketing, education and other commercial applications.
The launch represents a major expansion from earlier Wan models, particularly in video length and multimodal input. Wan3.0 can recommend an appropriate video duration based on a user’s prompt, extend existing videos and edit generated footage without requiring complete regeneration. Alibaba says the model also improves consistency across characters, products, locations and visual styles, although it acknowledges that audio texture and on-screen text rendering still require further refinement.
Alibaba Launches Wan3.0 AI Video Model
Wan3.0 is the newest generation of Alibaba’s Wan video-generation model family.
The model is designed to move beyond simple text-to-video generation by allowing users to provide different types of source material and transform them into finished video content.
Alibaba says Wan3.0 can generate up to 30 seconds of video in a single pass.
| Wan3.0 at a Glance | Details |
|---|---|
| Developer | Alibaba |
| Platform | Alibaba Cloud Model Studio |
| Maximum single-generation length | 30 seconds |
| Input types | Text, image, audio, video, documents |
| Document support | PDF, PPT, DOC, XLS, TXT, KEY, Pages, Numbers, Markdown |
| Maximum document size | 100 MB |
| Maximum document length | 50 pages |
| Output resolutions | 480P, 720P, 1080P |
| Starting API price | $0.05 per second |
| 720P price | $0.10 per second |
| 1080P price | $0.20 per second |
| Video editing | Supported |
| Availability | Alibaba Cloud Model Studio |
The combination of longer videos and broader inputs makes Wan3.0 significantly different from earlier versions of the model.
30-Second Video Generation Is the Biggest Upgrade
The most visible improvement is the ability to generate up to 30 seconds of video in a single generation.
Earlier Wan models were limited to shorter clips. Wan2.7-Video, for example, supported clips of around 15 seconds.
Wan3.0 doubles that maximum.
| Wan Generation | Maximum Single Generation |
|---|---|
| Earlier Wan models | Shorter clips |
| Wan2.7-Video | ~15 seconds |
| Wan3.0 | 30 seconds |
| Increase from Wan2.7-Video | 2X |
A longer generation window gives creators more room for continuous camera movements, uninterrupted shots and more complex storytelling.
From Shot to Story
Short AI clip
↓
Single visual moment
↓
Limited camera movement
↓
Wan3.0
↓
30-second sequence
↓
Multiple actions and camera movements
↓
More complete story
The longer duration does not eliminate the need for editing, but it can reduce the number of separate generations required to create a finished sequence.
Smart Duration Recommendation
Wan3.0 does not require users to generate the maximum 30-second video every time.
Alibaba has added a smart duration feature that recommends an appropriate length based on the user’s prompt.
A simple product animation may only require a few seconds.
A more complicated cinematic sequence may benefit from the full available duration.
This allows the system to match generation length with the complexity of the request.
How Smart Duration Works
User enters prompt
↓
Model analyzes requested action
↓
Determines appropriate duration
↓
Generates video
↓
User can extend it if necessary
The feature is intended to prevent short concepts from being unnecessarily padded while giving longer narratives enough space.
Video Extension Allows Longer Stories
Wan3.0 also supports video extension.
A creator can take an existing generated sequence and continue the story rather than starting a completely new generation.
This can be useful for longer narratives that exceed the initial 30-second generation limit.
Video Extension Workflow
Initial prompt
↓
30-second video
↓
Review sequence
↓
Extend video
↓
Continue story
↓
Additional scenes
↓
Longer narrative
This approach allows creators to build longer videos from multiple connected segments.
Wan3.0 Can Turn Documents Into Videos
One of the most important additions is document input.
Wan3.0 can accept documents and use their content as the basis for video generation.
Supported formats include DOC, XLS, PPT, PDF, TXT, KEY, Pages, Numbers and Markdown.
Alibaba says users can submit one file or link per request, with a maximum size of 100 MB and a maximum length of 50 pages.
| Document Capability | Limit |
|---|---|
| Files or links per request | 1 |
| Maximum file size | 100 MB |
| Maximum document length | 50 pages |
| Supported formats | DOC, XLS, PPT, PDF, TXT, KEY, Pages, Numbers, MD |
This feature expands Wan3.0 from a creative-generation system into a tool for transforming business information into visual content.
From PowerPoint to Video
A company could upload a product presentation and ask Wan3.0 to turn it into a promotional video.
A training department could provide a presentation and generate video course material.
A marketing team could transform a product document into a brand film.
Document-to-Video Examples
Product presentation
↓
Product information
↓
AI-generated product video
Training presentation
↓
Training material
↓
Video course
Business report
↓
Data and findings
↓
Narrated video briefing
Spreadsheet
↓
Data
↓
Animated charts and visual explanation
This “everything-to-video” approach is one of the defining features of Wan3.0.
Multimodal Inputs Expand Creative Control
Wan3.0 supports text, images, audio and video alongside documents.
That means users can provide multiple forms of reference material rather than relying entirely on a written prompt.
For example, a creator could provide a product image, a description and an audio reference.
The model can then use these inputs to produce a video that incorporates information from each source.
Multimodal Workflow
Text
+
Image
+
Audio
+
Video
+
Document
↓
Wan3.0
↓
Generated video
The approach gives creators greater control over the final output.
Character Consistency Gets a Major Upgrade
Consistency is one of the biggest challenges in AI-generated video.
Characters can change facial features, hairstyles, clothing or body proportions between frames or scenes.
Wan3.0 is designed to reduce this problem.
Alibaba says the model can maintain fine-grained consistency in facial features, hairstyles, clothing, body shape and accessories when reference material is provided.
Character Consistency
Reference character
↓
Face
↓
Hair
↓
Clothing
↓
Body shape
↓
Accessories
↓
Consistent video character
This is particularly important for storytelling because viewers quickly notice when an AI-generated character changes appearance.
Products and Props Can Stay Consistent
The model is also designed to maintain consistency in objects and products.
For commercial applications, this can be particularly important.
A product advertisement cannot show a smartphone changing its shape between shots or a vehicle logo becoming distorted.
Wan3.0 aims to maintain product details, hardware structures, logos and materials across generated sequences.
| Reference Element | Consistency Focus |
|---|---|
| Character | Face, hair, body, clothing |
| Product | Shape, hardware, logo |
| Props | Materials and appearance |
| Space | Layout and perspective |
| Style | Cinematic tone and visual identity |
| Audio | Voice and sound continuity |
These capabilities are intended to make AI-generated footage more usable in professional production environments.
Spatial Consistency Is Also Important
Wan3.0 is designed to maintain spatial relationships between people, objects and the camera.
This matters when characters move through a scene.
Without spatial consistency, AI-generated videos can produce physically confusing results, such as objects changing position or characters appearing to move unnaturally.
Alibaba says the model improves character blocking and camera perspective while maintaining spatial relationships.
Scene Consistency
Characters
+
Objects
+
Camera
+
Environment
↓
Stable spatial relationships
↓
More coherent video
This can make longer AI-generated sequences easier to use in professional editing.
More Realistic Human Faces
Alibaba is also emphasizing improvements to human faces.
The company says Wan3.0 is designed to produce more diverse and lifelike faces rather than repeatedly generating similar-looking AI characters.
The model also aims to produce more natural facial expressions.
Micro-expressions are linked with body movements and emotional context.
This is particularly useful for advertisements, short films, social videos and other content that depends on believable human performances.
Multilingual Voice Generation
Wan3.0 can also generate multilingual voice output.
This expands the model’s potential use for international advertising and educational content.
A company could potentially create versions of the same video for different markets without recording entirely separate productions.
Global Content Workflow
Create core video
↓
Select target language
↓
Generate multilingual voice
↓
Maintain visual sequence
↓
Produce localized version
The feature could reduce the cost and time involved in creating localized video content.
Video Editing Is Built Into Wan3.0
Wan3.0 does not only generate new footage.
It also carries forward video-editing capabilities introduced in Wan2.7.
Users can modify visual elements, plot details and dialogue without regenerating the entire video from scratch.
This changes the workflow from repeated generation to iterative editing.
Traditional AI Video Workflow
Generate
↓
Find problem
↓
Regenerate entire clip
↓
Find another problem
↓
Regenerate again
Wan3.0 Editing Workflow
Generate
↓
Identify problem
↓
Edit specific element
↓
Review
↓
Continue
This could significantly reduce wasted generation time and computing costs.
Wan3.0 Supports 480P, 720P and 1080P
Alibaba is offering three output-resolution options through its API.
The pricing is based on the number of seconds generated.
| Resolution | Price Per Second | 30-Second Generation |
|---|---|---|
| 480P | $0.05 | $1.50 |
| 720P | $0.10 | $3.00 |
| 1080P | $0.20 | $6.00 |
The pricing makes lower-resolution generation particularly useful for drafts.
Creators could generate an initial concept at 480P, make revisions and then produce the final version at 1080P.
Draft-to-Final Workflow
480P draft
↓
Review concept
↓
Edit
↓
Generate improved version
↓
1080P final
This could help control costs during the creative process.
API Pricing Starts at $0.05 Per Second
Wan3.0’s API starts at $0.05 per second for 480P output.
The price doubles to $0.10 per second at 720P and doubles again to $0.20 per second at 1080P.
For developers building video-generation applications, per-second pricing makes it easier to estimate costs based on expected output volume.
For example, generating 100 thirty-second 1080P videos would produce 3,000 seconds of video.
At $0.20 per second, the generation cost would be approximately $600 before any other applicable costs.
Advertising Could Be a Major Market
Advertising is one of the clearest commercial applications for Wan3.0.
Traditional advertisements can require actors, cameras, locations, lighting, editing and multiple production teams.
AI video generation can reduce the cost and time required to create certain types of promotional content.
Wan3.0’s ability to preserve product details could make it particularly useful for product-focused advertising.
AI Advertising Workflow
Product information
↓
Product images
↓
Creative concept
↓
Wan3.0
↓
Brand video
↓
Social media advertisement
↓
Multiple localized versions
The technology could be especially useful for companies producing large numbers of short marketing videos.
Social Media Creators Could Benefit
Short-form content creators are another target market.
Platforms such as YouTube, Instagram and TikTok require a constant supply of visual content.
Generating longer sequences from a single prompt could reduce the amount of manual production required.
Creators could also use document inputs to turn articles, reports or presentations into visual stories.
Film and Short-Drama Production
Wan3.0 is also being positioned for film and episodic content.
The combination of longer clips, character consistency and video extension can make AI-generated sequences more suitable for narrative storytelling.
The technology is not necessarily replacing traditional filmmaking.
Instead, it can be used for concept development, previsualization, background scenes, visual effects, storyboarding and lower-budget productions.
AI Film Workflow
Script
↓
Character references
↓
Scene references
↓
Wan3.0 generation
↓
Editing
↓
Visual refinement
↓
Final production
The longer generation window can make this process more practical than systems limited to very short clips.
Tourism Marketing Is Another Use Case
Alibaba also identifies tourism and destination marketing as potential applications.
A destination marketing organization could provide images and information about a location and generate promotional videos without organizing a large-scale physical shoot.
Natural landscapes, cultural landmarks and local food can be turned into visual promotional material.
Destination Marketing
Location images
+
Tourism information
+
Creative prompt
↓
AI-generated destination film
↓
Social media
↓
Travel campaigns
This could reduce production costs for smaller tourism organizations.
Education and Training Could Be Transformed
Documents-to-video capabilities may be particularly useful for education.
Teachers, training departments and businesses often work with PowerPoint presentations, PDFs and written manuals.
Wan3.0 can use those materials as source information for video generation.
A static training deck could therefore become a narrated instructional video.
Training Content
PDF/manual
↓
Wan3.0
↓
Narrated explanation
↓
Visual demonstrations
↓
Training video
This could help organizations create educational material without producing every video manually.
Data Visualization Becomes a Video Input
The model can also work with spreadsheets.
That creates a potential use case for financial reporting and business presentations.
A spreadsheet containing sales figures could be transformed into animated charts and visual explanations.
A business report could become a narrated executive briefing.
| Input | Potential Output |
|---|---|
| Spreadsheet | Animated data visualization |
| Financial report | Executive video briefing |
| Sales presentation | Product or sales video |
| Training deck | Instructional video |
| Product document | Marketing film |
This expands AI video generation into corporate communications.
Wan3.0 Could Help With Robotics and Autonomous Driving
Alibaba is also highlighting applications beyond media.
The company says Wan3.0 can generate realistic simulation videos that may be useful for training autonomous-driving and robotics systems.
Synthetic video data can help developers expose AI systems to different scenarios without physically recreating every situation.
Simulation Workflow
Real-world reference
↓
AI-generated scenario
↓
Different environmental conditions
↓
Synthetic video data
↓
Training
↓
Robotics or autonomous-driving model
This could become an important application for generative video as AI moves deeper into physical-world systems.
Wan3.0 Still Has Limitations
Despite the improvements, Alibaba acknowledges that Wan3.0 is not perfect.
The company specifically identifies audio texture and on-screen text rendering accuracy as areas that still need improvement.
This is important for commercial applications because text errors can make AI-generated advertisements, presentations and instructional videos unusable without manual correction.
Current Limitations
Audio texture
Still being refined
On-screen text
Accuracy can vary
Longer videos
May still require editing
Professional production
Human review remains important
The model therefore represents a significant step forward but is not a complete replacement for traditional production workflows.
Wan3.0 Has Evolved Through Eight Generations
Alibaba says the Wan family has gone through eight iterations from Wan1.0 to Wan3.0.
The progression reflects a broader shift in AI video generation.
Early models focused primarily on generating individual visual assets.
Newer models are increasingly designed to produce complete pieces of content and integrate into commercial workflows.
Wan Model Evolution
Wan1.0
Asset generation
↓
Later generations
Improved realism and control
↓
Wan2.x
Longer videos and editing
↓
Wan3.0
Everything-to-video
↓
30-second generation
↓
Professional workflows
Alibaba is positioning the latest generation as part of a longer-term effort to build models that can simulate and reproduce aspects of the physical world.
Wan3.0 vs Earlier Wan Models
The biggest differences can be summarized across five areas.
| Feature | Earlier Wan Models | Wan3.0 |
|---|---|---|
| Maximum single generation | Shorter clips | 30 seconds |
| Document input | Limited/absent | Yes |
| Multimodal input | Yes | Expanded |
| Character consistency | Improving | More precise |
| Video editing | Available in newer versions | Continued and expanded |
| Smart duration | Limited | Yes |
| Video extension | Available | Yes |
| Output resolution | Varies | 480P, 720P, 1080P |
The upgrade is therefore broader than simply increasing video length.
Wan3.0 Strengthens Alibaba’s AI Video Position
Alibaba is competing in an increasingly crowded AI video market.
Companies such as OpenAI, Google, ByteDance and other AI developers are investing heavily in systems that can generate realistic video.
The competition is moving beyond basic text-to-video generation.
Models increasingly need to understand reference images, maintain characters, follow camera instructions and generate longer sequences.
Alibaba’s document-to-video capability gives Wan3.0 another point of differentiation.
The AI Video Market Is Moving Toward Production Tools
The first wave of AI video tools focused heavily on novelty.
Users could generate impressive short clips from simple prompts.
The next stage is increasingly about production.
Professional users need:
- Consistency
- Editing
- Longer sequences
- Reference control
- Predictable output
- Multiple resolutions
- Lower costs
- Commercial APIs
Wan3.0 is designed around these requirements.
From AI Toy to Production Tool
Text prompt
↓
Interesting clip
↓
Reference control
↓
Longer video
↓
Editing
↓
Document input
↓
API integration
↓
Commercial production
This shift could determine which AI video platforms gain lasting business adoption.
Key Numbers at a Glance
30 seconds
Maximum single-generation video length
2X
Increase over Wan2.7-Video’s approximately 15-second maximum
$0.05
Price per second for 480P
$0.10
Price per second for 720P
$0.20
Price per second for 1080P
$1.50
Approximate cost of a 30-second 480P generation
$3.00
Approximate cost of a 30-second 720P generation
$6.00
Approximate cost of a 30-second 1080P generation
100 MB
Maximum supported document size
50 pages
Maximum supported document length
8
Wan model iterations from Wan1.0 to Wan3.0
What Wan3.0 Means for Creators
For creators, the biggest advantage is the ability to generate longer and more consistent sequences while using a broader range of reference material.
A creator can start with a written concept, provide images and audio references and generate a video.
If the result needs changes, the creator can edit elements rather than regenerate everything.
That could make AI video production more iterative and less wasteful.
What Wan3.0 Means for Businesses
For businesses, the document-to-video feature may be more important than the cinematic improvements.
Companies already have enormous amounts of content stored in PDFs, presentations, spreadsheets and documents.
Turning those assets into videos could create new ways to communicate with customers and employees.
Marketing teams could transform product documents into advertisements.
Human-resources teams could turn training materials into instructional videos.
Executives could convert reports into visual briefings.
What Wan3.0 Means for the AI Industry
Wan3.0 shows that the AI video market is moving toward multimodal systems capable of understanding and transforming information rather than simply responding to text prompts.
The ability to turn documents into video is particularly significant because it connects generative AI with existing business information.
If models become increasingly capable of understanding structured data, presentations and documents, video could become a new presentation layer for enterprise information.
What Developers Will Watch
Developers will likely focus on the model’s API performance, pricing and consistency.
The $0.05-per-second 480P starting price makes experimentation relatively inexpensive compared with high-resolution production.
The $0.20-per-second 1080P rate could also be attractive for certain commercial applications.
The bigger question will be whether the model can consistently deliver professional-quality results without requiring extensive manual correction.
Looking Ahead
Alibaba’s Wan3.0 marks a significant step in the evolution of AI video generation by combining 30-second single-pass video creation with text, image, audio, video and document inputs. The model can process files such as PDFs, PowerPoint presentations and spreadsheets, allowing users to transform existing business and creative material into dynamic video. Its API pricing starts at $0.05 per second for 480P, rising to $0.20 per second for 1080P, while features such as smart duration, video extension and in-place editing are designed to make the system more practical for production workflows.
The larger significance of Wan3.0 is its move from simple text-to-video generation toward an “everything-to-video” platform. Advertising, education, tourism, social media, filmmaking and enterprise communications could all benefit if the model delivers consistent results at scale. Alibaba still acknowledges limitations in audio texture and on-screen text rendering, but continued improvements in these areas could make AI-generated video increasingly viable as a commercial production tool rather than simply an experimental technology.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.