
Veo 3.1 and Wan 2.6 are both video models that generate sound with the picture, but they suit different jobs. Wan 2.6, from Alibaba, makes longer clips (up to 15 seconds), can split a prompt into several shots on its own and can take reference videos of a character, including their voice. Veo 3.1, from Google DeepMind, makes shorter clips that follow a detailed prompt closely, with start and end frame control and a 4K option.
This comparison explains what each difference means in practice. Specs were checked against Alibaba Cloud's and Google's own documentation on September 28, 2026.
The short answer
- Pick Wan 2.6 for 10–15 second clips, automatic multi-shot sequences, square or 4:3 formats, or when you want a character's appearance and voice carried in from a reference video.
- Pick Veo 3.1 for a single, precisely directed shot, a transition between two known frames, a 4K master, or dialogue written line by line in the prompt.
- Neither is downloadable. Wan 2.6 is a hosted model like Veo 3.1; Alibaba's public weights stop at Wan 2.2.
Veo 3.1 vs Wan 2.6 specs
| Veo 3.1 (Google DeepMind) | Wan 2.6 (Alibaba) | |
|---|---|---|
| Released | October 15, 2025 (4K and vertical upgrades January 13, 2026) | December 16, 2025 |
| Clip length | 4, 6 or 8 seconds in Google's API (8 s for 1080p and 4K) | 2–15 seconds (reference-to-video: 2–10 seconds) |
| Resolution | 720p, 1080p, 4K (Google describes 4K as upscaled) | 720p, 1080p |
| Aspect ratios | 16:9, 9:16 | 16:9, 9:16, 1:1, 4:3, 3:4 |
| Multi-shot | One shot by default; cuts only if you prompt for them | Built-in multi-shot mode |
| Starting points | Text; a start image; start + end frames; up to 3 reference images | Text; a first-frame image; up to 5 references (up to 3 videos, up to 5 images) |
| Audio | Generated with the video: dialogue, effects, ambience | Generated automatically, or you supply your own audio track |
| Weights | No public weights; hosted by Google | No public weights (Hugging Face lists up to Wan 2.2); hosted by Alibaba |
| Where to use it | Gemini app, Flow, Gemini API, Vertex AI; Veo Video | Alibaba Cloud Model Studio, the Wan website, Qwen app |
Clip length: 8 seconds vs up to 15
Wan 2.6's text-to-video model accepts any whole number of seconds from 2 to 15. That extra length is useful for a single continuous action that needs time to play out, or a short ad that should not be split across several generations. Its reference-to-video model tops out at 10 seconds.
Veo 3.1 works in shorter units: 4, 6 or 8 seconds in Google's API, with 1080p and 4K requiring the full 8. Google's tools can extend a clip (the API adds 7 seconds per step, up to 20 times, at 720p), but most longer Veo projects are built as several 8-second shots edited together. The upside is that each shot can be regenerated without touching the rest.
Multi-shot storytelling
Alibaba presents "intelligent multi-shot storytelling" as a headline feature of Wan 2.6. In the API you turn it on with a multi-shot setting (together with automatic prompt expansion), and the model splits the idea into several shots while keeping key details consistent. This is handy when you want a sequence — establishing shot, action, reaction — and do not want to plan each cut yourself.
Veo 3.1 defaults to one continuous shot. You can still ask for cuts in the prompt, and the model will often change framing within the 8 seconds, but you are directing that explicitly rather than handing the shot plan to the model.
What this means for you: if you prefer to let the model structure a short sequence, Wan 2.6 does more of that work. If you want control over every shot, Veo 3.1's one-shot-per-generation approach is easier to steer.
References: video vs images
Wan 2.6's reference-to-video model takes up to five references, of which up to three can be videos (each 1–30 seconds) and up to five can be images. Each reference should show a single character or object. According to Alibaba, the model extracts "character appearance and voice (if available)" from them, so a short clip of a person talking can carry both their look and their voice into the new scene.
Veo 3.1 uses still images. You can give it up to three reference images of a person, character or product to keep them consistent (Google calls this Ingredients to Video), or set a first frame and a last frame and let the model build the motion between them. Stills are easier to prepare; a reference video carries more information, such as how someone moves and sounds.
Audio
Both models generate sound with the video.
Veo 3.1 generates dialogue, sound effects and ambience, and you direct them in the prompt: spoken lines in quotes, named effects, a described ambience.
Wan 2.6 either generates matching background music and sound effects automatically or uses an audio file you upload. That second option is useful when the soundtrack already exists — a voiceover, a jingle — and the picture has to follow it.
Resolution and formats
Veo 3.1 offers 720p, 1080p and 4K in 16:9 or 9:16. Google describes its 4K as "state-of-the-art upscaling", so think of it as a sharper master for large screens.
Wan 2.6 offers 720p and 1080p, but in five shapes: 16:9, 9:16, 1:1, 4:3 and 3:4. If you need square posts or 4:3 frames straight out of the model, Wan 2.6 saves a crop.
Open source vs hosted
Earlier Wan releases were open-weight models, which is why Wan is often described as open source. That no longer applies to Wan 2.6. Alibaba's Wan-AI page on Hugging Face lists weights up to Wan 2.2, and Wan 2.6 is offered through Alibaba Cloud Model Studio, the Wan website and the Qwen app. If running a model on your own GPU is a requirement, Wan 2.2 is the relevant release, not Wan 2.6.
Veo 3.1 is also hosted only: you reach it through Google's products and API, or through platforms such as Veo Video.
Alibaba has also released newer Wan versions since Wan 2.6, including Wan 3.0, which its documentation lists with up to 30 seconds per generation. If you are choosing a Wan model today, compare against the current version too.
Which should you choose?
| Your project | Better fit | Why |
|---|---|---|
| A 10–15 second single take | Wan 2.6 | Up to 15 seconds per clip |
| A short sequence the model plans for you | Wan 2.6 | Built-in multi-shot mode |
| A character whose voice must carry over | Wan 2.6 | Reference videos pass on appearance and voice |
| A precise hero shot or product clip | Veo 3.1 | Close prompt following, start and end frames |
| A transition that lands on a known frame | Veo 3.1 | First and last frame control |
| A 4K master | Veo 3.1 | 720p, 1080p and 4K |
| Square or 4:3 social posts | Wan 2.6 | Five aspect ratios |
| Vertical Shorts, Reels or TikTok | Either | Both support 9:16 |
Trying Veo 3.1 on Veo Video
Wan 2.6 is not offered on Veo Video; Veo 3.1 is, in three tiers. Each generation is an 8-second clip with sound, in 16:9 or 9:16, at 720p, 1080p or 4K. You can start from text or a photo and add an end frame on any tier; reference images run on Veo 3.1 Fast, with up to three images per clip.
Draft the prompt on Veo 3.1 Lite at 720p, then generate the keeper on the full Veo 3.1 tier. New accounts get 30 free credits, enough for one Veo 3.1 Lite clip at 720p.
For more comparisons, see Veo 3.1 vs Seedance 2.0 and Veo 3.1 vs Kling AI.
