Wan 3.0 AI Video Generator
30 seconds in one take. Any input to video. Reference consistency. Native audio. Start from text, frames, a 10/5/5 pack, or a public document or webpage, then confirm 2–30 seconds and 480p–1080p before you generate.
- Alibaba · Wan 3.0 docs ↗
- Text, frames, pack, or page
- 2–30s · 480p–1080p
- From 18 credits
Create with Wan 3.0
Wan 3.0 Video Examples

Matcha powder pour
Prompt
Two hands open a matcha latte box on a pale pink tabletop, slide out a dark-green tray with a bamboo whisk, then dip fingers into a frosted glass jar of matcha powder and lift a clump so it falls back in a fine stream. Vertical framing, overhead and slightly angled, soft diffused studio light. Powder clings to the fingertips. Quiet room tone.
Key Features of Wan 3.0
A 2–30 second take from mixed inputs, with locked references and audio in the same pass—each row has a matching demo.

Native 30-second storytelling
Tell a complete beat in one take. Wan 3.0 holds opening, turn, and finish across a continuous 2–30 second clip—without stitching shorter takes.
How to Use Wan 3.0
- 1
Pick the simplest input that already holds the shot
Text for a new scene. A first frame plus an optional last frame to lock start and finish. A 10/5/5 pack when character, product, or place must stay assigned. Or one public document or webpage URL. First/last-frame cannot be combined with a pack or page. Skip extra files that do not have a job. Prefill from a demo if you want—nothing runs until you generate.
- 2
Write how the scene unfolds
Subject, action, setting, camera, then sound. For 2–30 seconds, cover an opening, a turn, and a finish—not a pile of unrelated events. Address Image 1, Video 1, or Audio 1 and say what each file provides. Name who speaks and the room when you need dialogue.
- 3
Generate, review, then iterate
Choose 2–30 seconds and 480p–1080p to match the brief, not the longest setting by default. Confirm the credit estimate, then generate. Watch the middle and the end—not just the first second—for identity, logos, dialogue, on-screen text, and lip sync. Change one or two things before you iterate.
Wan 3.0 Pricing
From 18 credits per 2-second 480p clip, up to 1,050 for 30 seconds at 1080p. Duration and 480p–1080p change the quote. Confirm the exact reservation before you generate.
How Wan 3.0 compares to other models
Pick by the job: one-take length, reference-pack size, document or webpage input, or an instruction-led edit. Specs below are the live MixVio routes.
| Attribute | This modelWan 3.0 | ByteDanceSeedance 2.5 → | MiniMaxMiniMax H3 → | KuaishouKling 3.0 → |
|---|---|---|---|---|
| Best for | 2–30s one-takes, 10/5/5 reference packs, and document or webpage input | 4–30s one-takes, packs up to 50 files, and clay-render blocking | Multimodal reference, first/last frames, and instruction-led edits | 5/10/15s cinematic shots from text or an owned frame |
| Inputs | Text · Up to 10 images · 5 videos · 5 audio | Text · Up to 30 images · 10 videos · 10 audio | Text · First/last frame · Up to 9 images · 3 videos · 3 audio | Text · Image + end frame |
| Output | 480p–1080p video + audio | 480p–1080p video + audio | 768p–2K video + audio | 720p–4K video + audio |
| Duration | 2–30s | 4–30s | 4–15s | 5–15s |
| Native audio | Yes | Yes | Yes | Yes |
| Credits | 18–1,050 credits | 156–7140 credits | 112–1406 credits | 110–330 credits |
Stay on Wan 3.0 for a 2–30 second take, 480p–1080p, a 10/5/5 pack, first/last frames, or a public document or webpage. Switch to Seedance 2.5 for packs up to 50 files, MiniMax H3 for 4–15s instruction-led edits at 768p or 2K, or Kling 3.0 for 5/10/15s cinematic shots.
Wan 3.0 FAQs
What is Wan 3.0?
Wan 3.0 is Alibaba Tongyi Lab’s video model for one continuous take. On MixVio it turns text, a first frame with an optional last frame, a 10/5/5 reference pack, or one public document or webpage into a 2–30 second clip at 480p, 720p, or 1080p, with audio in the same pass. Alibaba Cloud Model Studio ↗
What is Wan 3.0 best for?
Use it for a 2–30 second one-take: a short story with an opening, turn, and finish; a product or character lock from owned references; an approved first and last frame; or a public document or webpage that already holds the brief. It is not the MixVio route for 4K output, in-place video edit, or official continuation.
What inputs does Wan 3.0 accept?
Use text for a new scene. Use a first frame plus an optional last frame when the start or finish is already approved. Use a 10/5/5 pack—up to 10 images, 5 videos (1–15s each, 15s combined), and 5 audio clips (1–15s each, 15s combined)—when identity, motion, or timing must stay assigned. You can also paste one public https document URL or one public webpage URL. First/last-frame mode cannot be combined with the pack, document, or webpage. Audio should not be the only reference.
How do I keep a face, product, or place consistent?
Give Wan the simplest input that already holds that identity—an approved first frame, an optional last frame, or a small pack. In the prompt, address Image 1, Video 1, or Audio 1 by job, and keep the same face, wardrobe, product, and setting. Do not rewrite the identity between drafts. Fine details can still drift on a long take or many scene changes—review faces, logos, and labels before publishing.
Can I extend or edit an existing Wan 3.0 video?
You can add an owned clip to the 10/5/5 pack (1–15s each, 15s combined), address Video 1, and describe what should happen next. MixVio still returns one new 2–30 second clip. Official video continuation, interval edit, and in-place timeline tools are not on the public MixVio Wan 3.0 routes.
How long can a Wan 3.0 video be?
Each MixVio run is one 2–30 second clip, default 5 seconds. Pick a length that fits the opening, turn, and finish—do not default to 30 seconds. MixVio keeps an explicit picker so credits are shown first. If you add reference video, input seconds plus output seconds cannot exceed 30.
Does Wan 3.0 generate audio with the video?
Yes. Dialogue, ambience, and score generate in the same pass, and the credit quote does not change. Name who speaks and the room in the brief, then review lip sync, the mix, and on-screen text before publishing—Alibaba notes those are still improving.
How many credits does Wan 3.0 use?
Wan 3.0 is not available on the Free plan. Paid plans reserve from 18 credits. Typical anchors: 18 credits for 2 seconds at 480p, 175 for 5 seconds at 1080p, 540 for 30 seconds at 720p, and 1,050 for 30 seconds at 1080p. Duration and 480p–1080p change the quote. The generator shows the exact reservation before you submit. Failed runs cost no credits.
How does Wan 3.0 compare with Seedance 2.5, MiniMax H3, and Kling 3.0?
Stay on Wan 3.0 for a 2–30 second take, 480p–1080p, a 10/5/5 pack, first/last frames, or a public document or webpage. Seedance 2.5 also reaches 30 seconds and accepts a larger 30/10/10 pack (50 files). MiniMax H3 is 4–15 seconds at 768p or 2K with instruction-led edits. Kling 3.0 Standard is 5/10/15 second cinematic shots. Compare live credits, inputs, duration, and audio before generating.
Is Wan 3.0 open source? Does it output 4K?
No. Official and Provider docs list 480p, 720p, and 1080p only. Alibaba’s published open weights stop at Wan 2.2. Pages that advertise Wan 3.0 4K or downloadable weights are not describing the live API MixVio calls.
Can I use Wan 3.0 videos commercially?
Paid-plan outputs may be used commercially, subject to the Terms of Service and current Alibaba / Provider terms. You must have rights to every prompt, reference, likeness, brand asset, voice, document, and webpage. Review identity, logos, dialogue, lip sync, consent, and audio before publishing. MixVio does not guarantee exclusivity or trademark clearance.
Start a 30-second take with Wan 3.0.
Private. Credits shown first.
Start from text, frames, a reference pack, or a public page, then generate with Wan 3.0 selected.
Create with Wan 3.0 ↑