Alibaba has launched Wan3.0, the latest generation of its Wan AI video-generation model, with the ability to create videos of up to 30 seconds in a single generation. The model is now available through Alibaba Cloud Model Studio and expands the types of material that can be converted into video, supporting text, images, audio, video and, for the first time in the Wan family, documents such as PDFs, PowerPoint presentations, spreadsheets and other files. Alibaba is positioning the model as a tool for professional creative work, advertising, marketing, education and other commercial applications.

The launch represents a major expansion from earlier Wan models, particularly in video length and multimodal input. Wan3.0 can recommend an appropriate video duration based on a user’s prompt, extend existing videos and edit generated footage without requiring complete regeneration. Alibaba says the model also improves consistency across characters, products, locations and visual styles, although it acknowledges that audio texture and on-screen text rendering still require further refinement.

Alibaba Launches Wan3.0 AI Video Model

Wan3.0 is the newest generation of Alibaba’s Wan video-generation model family.

The model is designed to move beyond simple text-to-video generation by allowing users to provide different types of source material and transform them into finished video content.

Alibaba says Wan3.0 can generate up to 30 seconds of video in a single pass.

Wan3.0 at a GlanceDetails
DeveloperAlibaba
PlatformAlibaba Cloud Model Studio
Maximum single-generation length30 seconds
Input typesText, image, audio, video, documents
Document supportPDF, PPT, DOC, XLS, TXT, KEY, Pages, Numbers, Markdown
Maximum document size100 MB
Maximum document length50 pages
Output resolutions480P, 720P, 1080P
Starting API price$0.05 per second
720P price$0.10 per second
1080P price$0.20 per second
Video editingSupported
AvailabilityAlibaba Cloud Model Studio

The combination of longer videos and broader inputs makes Wan3.0 significantly different from earlier versions of the model.

30-Second Video Generation Is the Biggest Upgrade

The most visible improvement is the ability to generate up to 30 seconds of video in a single generation.

Earlier Wan models were limited to shorter clips. Wan2.7-Video, for example, supported clips of around 15 seconds.

Wan3.0 doubles that maximum.

Wan GenerationMaximum Single Generation
Earlier Wan modelsShorter clips
Wan2.7-Video~15 seconds
Wan3.030 seconds
Increase from Wan2.7-Video2X

A longer generation window gives creators more room for continuous camera movements, uninterrupted shots and more complex storytelling.

From Shot to Story

Short AI clip

Single visual moment

Limited camera movement

Wan3.0

30-second sequence

Multiple actions and camera movements

More complete story

The longer duration does not eliminate the need for editing, but it can reduce the number of separate generations required to create a finished sequence.

Smart Duration Recommendation

Wan3.0 does not require users to generate the maximum 30-second video every time.

Alibaba has added a smart duration feature that recommends an appropriate length based on the user’s prompt.

A simple product animation may only require a few seconds.

A more complicated cinematic sequence may benefit from the full available duration.

This allows the system to match generation length with the complexity of the request.

How Smart Duration Works

User enters prompt

Model analyzes requested action

Determines appropriate duration

Generates video

User can extend it if necessary

The feature is intended to prevent short concepts from being unnecessarily padded while giving longer narratives enough space.

Video Extension Allows Longer Stories

Wan3.0 also supports video extension.

A creator can take an existing generated sequence and continue the story rather than starting a completely new generation.

This can be useful for longer narratives that exceed the initial 30-second generation limit.

Video Extension Workflow

Initial prompt

30-second video

Review sequence

Extend video

Continue story

Additional scenes

Longer narrative

This approach allows creators to build longer videos from multiple connected segments.

Wan3.0 Can Turn Documents Into Videos

One of the most important additions is document input.

Wan3.0 can accept documents and use their content as the basis for video generation.

Supported formats include DOC, XLS, PPT, PDF, TXT, KEY, Pages, Numbers and Markdown.

Alibaba says users can submit one file or link per request, with a maximum size of 100 MB and a maximum length of 50 pages.

Document CapabilityLimit
Files or links per request1
Maximum file size100 MB
Maximum document length50 pages
Supported formatsDOC, XLS, PPT, PDF, TXT, KEY, Pages, Numbers, MD

This feature expands Wan3.0 from a creative-generation system into a tool for transforming business information into visual content.

From PowerPoint to Video

A company could upload a product presentation and ask Wan3.0 to turn it into a promotional video.

A training department could provide a presentation and generate video course material.

A marketing team could transform a product document into a brand film.

Document-to-Video Examples

Product presentation

Product information

AI-generated product video

Training presentation

Training material

Video course

Business report

Data and findings

Narrated video briefing

Spreadsheet

Data

Animated charts and visual explanation

This “everything-to-video” approach is one of the defining features of Wan3.0.

Multimodal Inputs Expand Creative Control

Wan3.0 supports text, images, audio and video alongside documents.

That means users can provide multiple forms of reference material rather than relying entirely on a written prompt.

For example, a creator could provide a product image, a description and an audio reference.

The model can then use these inputs to produce a video that incorporates information from each source.

Multimodal Workflow

Text

+

Image

+

Audio

+

Video

+

Document

Wan3.0

Generated video

The approach gives creators greater control over the final output.

Character Consistency Gets a Major Upgrade

Consistency is one of the biggest challenges in AI-generated video.

Characters can change facial features, hairstyles, clothing or body proportions between frames or scenes.

Wan3.0 is designed to reduce this problem.

Alibaba says the model can maintain fine-grained consistency in facial features, hairstyles, clothing, body shape and accessories when reference material is provided.

Character Consistency

Reference character

Face

Hair

Clothing

Body shape

Accessories

Consistent video character

This is particularly important for storytelling because viewers quickly notice when an AI-generated character changes appearance.

Products and Props Can Stay Consistent

The model is also designed to maintain consistency in objects and products.

For commercial applications, this can be particularly important.

A product advertisement cannot show a smartphone changing its shape between shots or a vehicle logo becoming distorted.

Wan3.0 aims to maintain product details, hardware structures, logos and materials across generated sequences.

Reference ElementConsistency Focus
CharacterFace, hair, body, clothing
ProductShape, hardware, logo
PropsMaterials and appearance
SpaceLayout and perspective
StyleCinematic tone and visual identity
AudioVoice and sound continuity

These capabilities are intended to make AI-generated footage more usable in professional production environments.

Spatial Consistency Is Also Important

Wan3.0 is designed to maintain spatial relationships between people, objects and the camera.

This matters when characters move through a scene.

Without spatial consistency, AI-generated videos can produce physically confusing results, such as objects changing position or characters appearing to move unnaturally.

Alibaba says the model improves character blocking and camera perspective while maintaining spatial relationships.

Scene Consistency

Characters

+

Objects

+

Camera

+

Environment

Stable spatial relationships

More coherent video

This can make longer AI-generated sequences easier to use in professional editing.

More Realistic Human Faces

Alibaba is also emphasizing improvements to human faces.

The company says Wan3.0 is designed to produce more diverse and lifelike faces rather than repeatedly generating similar-looking AI characters.

The model also aims to produce more natural facial expressions.

Micro-expressions are linked with body movements and emotional context.

This is particularly useful for advertisements, short films, social videos and other content that depends on believable human performances.

Multilingual Voice Generation

Wan3.0 can also generate multilingual voice output.

This expands the model’s potential use for international advertising and educational content.

A company could potentially create versions of the same video for different markets without recording entirely separate productions.

Global Content Workflow

Create core video

Select target language

Generate multilingual voice

Maintain visual sequence

Produce localized version

The feature could reduce the cost and time involved in creating localized video content.

Video Editing Is Built Into Wan3.0

Wan3.0 does not only generate new footage.

It also carries forward video-editing capabilities introduced in Wan2.7.

Users can modify visual elements, plot details and dialogue without regenerating the entire video from scratch.

This changes the workflow from repeated generation to iterative editing.

Traditional AI Video Workflow

Generate

Find problem

Regenerate entire clip

Find another problem

Regenerate again

Wan3.0 Editing Workflow

Generate

Identify problem

Edit specific element

Review

Continue

This could significantly reduce wasted generation time and computing costs.

Wan3.0 Supports 480P, 720P and 1080P

Alibaba is offering three output-resolution options through its API.

The pricing is based on the number of seconds generated.

ResolutionPrice Per Second30-Second Generation
480P$0.05$1.50
720P$0.10$3.00
1080P$0.20$6.00

The pricing makes lower-resolution generation particularly useful for drafts.

Creators could generate an initial concept at 480P, make revisions and then produce the final version at 1080P.

Draft-to-Final Workflow

480P draft

Review concept

Edit

Generate improved version

1080P final

This could help control costs during the creative process.

API Pricing Starts at $0.05 Per Second

Wan3.0’s API starts at $0.05 per second for 480P output.

The price doubles to $0.10 per second at 720P and doubles again to $0.20 per second at 1080P.

For developers building video-generation applications, per-second pricing makes it easier to estimate costs based on expected output volume.

For example, generating 100 thirty-second 1080P videos would produce 3,000 seconds of video.

At $0.20 per second, the generation cost would be approximately $600 before any other applicable costs.

Advertising Could Be a Major Market

Advertising is one of the clearest commercial applications for Wan3.0.

Traditional advertisements can require actors, cameras, locations, lighting, editing and multiple production teams.

AI video generation can reduce the cost and time required to create certain types of promotional content.

Wan3.0’s ability to preserve product details could make it particularly useful for product-focused advertising.

AI Advertising Workflow

Product information

Product images

Creative concept

Wan3.0

Brand video

Social media advertisement

Multiple localized versions

The technology could be especially useful for companies producing large numbers of short marketing videos.

Social Media Creators Could Benefit

Short-form content creators are another target market.

Platforms such as YouTube, Instagram and TikTok require a constant supply of visual content.

Generating longer sequences from a single prompt could reduce the amount of manual production required.

Creators could also use document inputs to turn articles, reports or presentations into visual stories.

Film and Short-Drama Production

Wan3.0 is also being positioned for film and episodic content.

The combination of longer clips, character consistency and video extension can make AI-generated sequences more suitable for narrative storytelling.

The technology is not necessarily replacing traditional filmmaking.

Instead, it can be used for concept development, previsualization, background scenes, visual effects, storyboarding and lower-budget productions.

AI Film Workflow

Script

Character references

Scene references

Wan3.0 generation

Editing

Visual refinement

Final production

The longer generation window can make this process more practical than systems limited to very short clips.

Tourism Marketing Is Another Use Case

Alibaba also identifies tourism and destination marketing as potential applications.

A destination marketing organization could provide images and information about a location and generate promotional videos without organizing a large-scale physical shoot.

Natural landscapes, cultural landmarks and local food can be turned into visual promotional material.

Destination Marketing

Location images

+

Tourism information

+

Creative prompt

AI-generated destination film

Social media

Travel campaigns

This could reduce production costs for smaller tourism organizations.

Education and Training Could Be Transformed

Documents-to-video capabilities may be particularly useful for education.

Teachers, training departments and businesses often work with PowerPoint presentations, PDFs and written manuals.

Wan3.0 can use those materials as source information for video generation.

A static training deck could therefore become a narrated instructional video.

Training Content

PDF/manual

Wan3.0

Narrated explanation

Visual demonstrations

Training video

This could help organizations create educational material without producing every video manually.

Data Visualization Becomes a Video Input

The model can also work with spreadsheets.

That creates a potential use case for financial reporting and business presentations.

A spreadsheet containing sales figures could be transformed into animated charts and visual explanations.

A business report could become a narrated executive briefing.

InputPotential Output
SpreadsheetAnimated data visualization
Financial reportExecutive video briefing
Sales presentationProduct or sales video
Training deckInstructional video
Product documentMarketing film

This expands AI video generation into corporate communications.

Wan3.0 Could Help With Robotics and Autonomous Driving

Alibaba is also highlighting applications beyond media.

The company says Wan3.0 can generate realistic simulation videos that may be useful for training autonomous-driving and robotics systems.

Synthetic video data can help developers expose AI systems to different scenarios without physically recreating every situation.

Simulation Workflow

Real-world reference

AI-generated scenario

Different environmental conditions

Synthetic video data

Training

Robotics or autonomous-driving model

This could become an important application for generative video as AI moves deeper into physical-world systems.

Wan3.0 Still Has Limitations

Despite the improvements, Alibaba acknowledges that Wan3.0 is not perfect.

The company specifically identifies audio texture and on-screen text rendering accuracy as areas that still need improvement.

This is important for commercial applications because text errors can make AI-generated advertisements, presentations and instructional videos unusable without manual correction.

Current Limitations

Audio texture

Still being refined

On-screen text

Accuracy can vary

Longer videos

May still require editing

Professional production

Human review remains important

The model therefore represents a significant step forward but is not a complete replacement for traditional production workflows.

Wan3.0 Has Evolved Through Eight Generations

Alibaba says the Wan family has gone through eight iterations from Wan1.0 to Wan3.0.

The progression reflects a broader shift in AI video generation.

Early models focused primarily on generating individual visual assets.

Newer models are increasingly designed to produce complete pieces of content and integrate into commercial workflows.

Wan Model Evolution

Wan1.0

Asset generation

Later generations

Improved realism and control

Wan2.x

Longer videos and editing

Wan3.0

Everything-to-video

30-second generation

Professional workflows

Alibaba is positioning the latest generation as part of a longer-term effort to build models that can simulate and reproduce aspects of the physical world.

Wan3.0 vs Earlier Wan Models

The biggest differences can be summarized across five areas.

FeatureEarlier Wan ModelsWan3.0
Maximum single generationShorter clips30 seconds
Document inputLimited/absentYes
Multimodal inputYesExpanded
Character consistencyImprovingMore precise
Video editingAvailable in newer versionsContinued and expanded
Smart durationLimitedYes
Video extensionAvailableYes
Output resolutionVaries480P, 720P, 1080P

The upgrade is therefore broader than simply increasing video length.

Wan3.0 Strengthens Alibaba’s AI Video Position

Alibaba is competing in an increasingly crowded AI video market.

Companies such as OpenAI, Google, ByteDance and other AI developers are investing heavily in systems that can generate realistic video.

The competition is moving beyond basic text-to-video generation.

Models increasingly need to understand reference images, maintain characters, follow camera instructions and generate longer sequences.

Alibaba’s document-to-video capability gives Wan3.0 another point of differentiation.

The AI Video Market Is Moving Toward Production Tools

The first wave of AI video tools focused heavily on novelty.

Users could generate impressive short clips from simple prompts.

The next stage is increasingly about production.

Professional users need:

  • Consistency
  • Editing
  • Longer sequences
  • Reference control
  • Predictable output
  • Multiple resolutions
  • Lower costs
  • Commercial APIs

Wan3.0 is designed around these requirements.

From AI Toy to Production Tool

Text prompt

Interesting clip

Reference control

Longer video

Editing

Document input

API integration

Commercial production

This shift could determine which AI video platforms gain lasting business adoption.

Key Numbers at a Glance

30 seconds

Maximum single-generation video length

2X

Increase over Wan2.7-Video’s approximately 15-second maximum

$0.05

Price per second for 480P

$0.10

Price per second for 720P

$0.20

Price per second for 1080P

$1.50

Approximate cost of a 30-second 480P generation

$3.00

Approximate cost of a 30-second 720P generation

$6.00

Approximate cost of a 30-second 1080P generation

100 MB

Maximum supported document size

50 pages

Maximum supported document length

8

Wan model iterations from Wan1.0 to Wan3.0

What Wan3.0 Means for Creators

For creators, the biggest advantage is the ability to generate longer and more consistent sequences while using a broader range of reference material.

A creator can start with a written concept, provide images and audio references and generate a video.

If the result needs changes, the creator can edit elements rather than regenerate everything.

That could make AI video production more iterative and less wasteful.

What Wan3.0 Means for Businesses

For businesses, the document-to-video feature may be more important than the cinematic improvements.

Companies already have enormous amounts of content stored in PDFs, presentations, spreadsheets and documents.

Turning those assets into videos could create new ways to communicate with customers and employees.

Marketing teams could transform product documents into advertisements.

Human-resources teams could turn training materials into instructional videos.

Executives could convert reports into visual briefings.

What Wan3.0 Means for the AI Industry

Wan3.0 shows that the AI video market is moving toward multimodal systems capable of understanding and transforming information rather than simply responding to text prompts.

The ability to turn documents into video is particularly significant because it connects generative AI with existing business information.

If models become increasingly capable of understanding structured data, presentations and documents, video could become a new presentation layer for enterprise information.

What Developers Will Watch

Developers will likely focus on the model’s API performance, pricing and consistency.

The $0.05-per-second 480P starting price makes experimentation relatively inexpensive compared with high-resolution production.

The $0.20-per-second 1080P rate could also be attractive for certain commercial applications.

The bigger question will be whether the model can consistently deliver professional-quality results without requiring extensive manual correction.

Looking Ahead

Alibaba’s Wan3.0 marks a significant step in the evolution of AI video generation by combining 30-second single-pass video creation with text, image, audio, video and document inputs. The model can process files such as PDFs, PowerPoint presentations and spreadsheets, allowing users to transform existing business and creative material into dynamic video. Its API pricing starts at $0.05 per second for 480P, rising to $0.20 per second for 1080P, while features such as smart duration, video extension and in-place editing are designed to make the system more practical for production workflows.

The larger significance of Wan3.0 is its move from simple text-to-video generation toward an “everything-to-video” platform. Advertising, education, tourism, social media, filmmaking and enterprise communications could all benefit if the model delivers consistent results at scale. Alibaba still acknowledges limitations in audio texture and on-screen text rendering, but continued improvements in these areas could make AI-generated video increasingly viable as a commercial production tool rather than simply an experimental technology.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.