Wan 3.0 AI Video Generator: 30-Second Clips With Sound

Alibaba’s newest Wan model turns a prompt, a first and last frame, or a stack of reference images, clips, and audio into one video of up to 30 seconds, in up to 1080p, with dialogue, music, and sound effects generated alongside the picture.

Need ideas?
Example

Generate a 30-second, 16:9 urban chase short in a stylish motion-graphics illustration style: flat graphic shapes, refined 2.5D layering, restrained 3D volume, not live action. Image1 is the reference for the neon city, roads, and overall texture; Image2 is the only reference for the man; Image3 is the only reference for the woman. Keep both faces, hairstyles, outfits, and colors consistent. Both are martial artists fighting over one glowing ball with a white core and a purple-pink-blue halo. 0–4 s: the camera dives from the neon skyline to street level as the woman sprints into frame clutching the ball. 4–8 s: the man cuts in from a side street to intercept her; she redirects his grab and slips past. 8–24 s: the chase runs through crowds and alleys, then both grab bicycles and race through night traffic, with handheld tracking, whip pans, snap zooms, and low angles. 24–30 s: she cuts between a braking car and a bus, escapes, and glances back with a proud smile as the camera rises over the city. Only one glowing ball, always the same size.

Reference example. It is not a result generated from your upload.

Your settings are ready

This is an interface preview. Your prompt has not been sent to an AI model, no video has been generated, and no credits have been used.

Describe your Wan 3.0 video
—
Wan model
Wan 3.0
Resolution
720P
Duration
5
Aspect ratio
Auto
Generate audio
On
Output count
1

What Is Wan 3.0?

Wan 3.0 (also written Wan3.0 or Wan 3) is the latest video generation model from Alibaba’s Wan team. It opened as a public beta on August 6, 2026 and launched widely on August 24, 2026 through Alibaba Cloud Model Studio. Alibaba calls it an all-in-one model: the same model handles text to video, image to video with a first and last frame, and reference to video, and it creates clips of up to 30 seconds with a native audio track, about twice as long as earlier Wan models.

Alibaba Wan 3.0 comes in two tiers, and both are in the workspace above: Wan 3.0, the standard model, and Wan 3.0 Prime, a faster version with the same inputs and settings at a higher cost. Where Wan 2.2 is built for quick 5-second clips, Wan 3.0 is built for longer, reference-heavy scenes in which a character, a product, or a location has to look the same from the first second to the last.

Wan 3.0 Specs: Wan 3.0 vs Wan 3.0 Prime

What you can set when you generate with each Wan 3.0 tier on EzImgMaker.

Wan 3.0Wan 3.0 Prime
Best forMost videos, at the lower costFaster results when you are short on time
Text to videoYesYes
Image to videoFirst frame, or first + last frameFirst frame, or first + last frame
Reference to videoUp to 10 images, 5 videos, and 5 audio clipsUp to 10 images, 5 videos, and 5 audio clips
Video length2–30 s2–30 s
Resolution480P, 720P, 1080P480P, 720P, 1080P
Aspect ratioAuto, 16:9, 9:16, 1:1, 4:3, 3:4Auto, 16:9, 9:16, 1:1, 4:3, 3:4
Native audioYes (on or off)Yes (on or off)
PromptUp to 20,000 characters, English or ChineseUp to 20,000 characters, English or Chinese

First/last frames and reference files are separate modes and cannot be mixed in one generation. Reference videos run 1–15 s each and 15 s in total, and the reference clip length plus the output length must stay within 30 seconds. Reference audio works best paired with an image or a video. This workspace takes up to 3 reference images, 1 video, and 1 audio clip per generation.

Wan 3.0 vs Wan 2.2

Which Wan model to pick on EzImgMaker.

Wan 3.0Wan 2.2 (A14B Turbo)
Best forLonger scenes with sound and consistent charactersQuick, low-cost 5-second drafts and talking portraits
Video length2–30 s, you chooseAbout 5 s, fixed
Resolution480P, 720P, 1080P480p or 720p
Aspect ratioAuto plus five ratios16:9 or 9:16
Native audioYesNo (speech-to-video lip-syncs to audio you upload)
Image inputFirst frame, or first + last frameOne starting image
ReferencesImages, videos, and audioNo
Open weightsNoYes (Apache 2.0)

Wan 2.2 has its own page with text-to-video, image-to-video, and speech-to-video modes.

Wan 3.0 Video Examples

Placeholder clips with the prompts behind them, trimmed to 8 seconds. Copy one and swap in your own subject.

How to Use the Wan 3.0 AI Video Generator

Three steps in your browser. No local install or GPU needed.

01

Pick an Input Mode and Write the Prompt

Use Text to video for a scene from words, Image to video to start on a first frame (and optionally end on a last frame), or Reference to video to add images, a video, and audio you call Image1, Video1, and Audio1.

02

Choose Model and Settings

Select Wan 3.0 or the faster Wan 3.0 Prime, then 480P, 720P, or 1080P, a length from 2 to 30 seconds, an aspect ratio or Auto, and whether to generate audio.

03

Generate and Refine

Create the Wan 3.0 video, check motion, timing, and consistency, then download it or change one beat of the prompt and run it again.

What Wan 3.0 Can Do

Up to 30 Seconds in One Generation

Wan 3.0 creates anything from a 2-second loop to a 30-second scene in a single pass, so a story can have a setup, a turn, and a payoff without stitching clips together. Write time codes such as 0–8 s and 8–16 s to pace each beat, and draft in 480P before you render the keeper in 1080P.

Wan 3.0 Reference to Video: Image1, Video1, Audio1

Upload reference images for faces, outfits, products, or places, a reference video for motion and camera work, and reference audio for a voice or a song, then refer to them as Image1, Video1, and Audio1 in the prompt. Giving each file a clear role keeps characters and props consistent across the whole clip.

Wan 3.0 Image to Video With First and Last Frames

Start the video on an exact first frame, such as a product shot, a storyboard panel, or a photo you own, and add a last frame when you know where the shot has to land. Wan 3.0 fills in the motion between them, and Auto aspect ratio follows the shape of your image.

Native Audio: Dialogue, Music, and Effects

With audio on, Wan 3.0 generates speech, ambience, music, and sound effects together with the picture, so a line of dialogue or a drum hit lands on the right frame. Write who says what and in which language, or switch audio off for silent B-roll you will score later.

What People Make With Alibaba Wan 3.0

Alibaba reports Wan 3.0 being used for these kinds of projects since its public beta.

Short Dramas and Film Previz

Longer scenes with recurring characters, dialogue, and planned camera moves.

Advertising and Marketing

Product films and social ads where the product and logo must stay identical in every shot.

Travel and Tourism Promotion

Scenic flyovers, city walks, and destination teasers with ambient sound.

Music Videos

Performances and dance clips driven by a reference song or voice.

Wan 3.0 Prompt Tips

A simple formula: subject + setting + action in order + camera + style + sound + what must not change.

Time-Code Long Clips

Split 15–30 second videos into beats like 0–8 s, 8–16 s, 16–30 s, and give each beat one main action.

Give Every Reference a Role

Write "Image1 is her face and outfit, Video1 is the camera move only, Audio1 is the song" so files are not blended by accident.

Write the Soundtrack

Name the speaker and language for each line, then list ambience, music, and key sound effects.

Say What Must Stay

End with constraints such as "keep the bottle shape, label, and color unchanged" or "only one glowing ball".

Wan 3.0 FAQ

Common questions about Alibaba Wan 3.0.

Make Your First Wan 3.0 Video

Write a 30-second story or upload your references, and let Wan 3.0 generate the picture and the sound together.

Wan, Alibaba, and Alibaba Cloud are trademarks of Alibaba Group Holding Limited or its affiliates. EzImgMaker is an independent service and is not affiliated with, endorsed by, or sponsored by Alibaba, Alibaba Cloud, or the Wan team. Model names are used only to describe compatibility.