
The answer has changed since this comparison was first written. OpenAI announced on March 24, 2026 that it was discontinuing Sora. The Sora app and website closed on April 26, 2026, and the Sora 2 models and Videos API were shut down on September 24, 2026. So today the practical question is not "Veo 3.1 or Sora 2?" but "what did Sora 2 do, and how much of it does Veo 3.1 cover?"
Short version: Veo 3.1 matches Sora 2's core — video and sound generated together from text or an image — and offers more control over individual shots: a start frame, an optional end frame, up to three reference images, and output up to 4K. What it doesn't replace is Sora's longer single clips (15 to 25 seconds) and the Cameo feature for putting yourself into a scene. For the full shutdown timeline, see Sora 2 is shutting down.
Veo 3.1 vs Sora 2: specs side by side
Veo 3.1 figures come from Google's Gemini API documentation; Sora 2 figures come from OpenAI's model pages and its October 2025 product update. All checked September 28, 2026.
| Veo 3.1 (Google DeepMind) | Sora 2 (OpenAI) | |
|---|---|---|
| Released | October 15, 2025 (updated January 2026) | September 30, 2025 |
| Status | Available: Gemini app, Flow, Gemini API, Vertex AI | App closed April 26, 2026; API closed September 24, 2026 |
| Native audio | Yes — effects, ambience, dialogue; always on | Yes — dialogue and sound effects |
| Clip length | 4, 6 or 8 seconds | Up to 15 s in the app for all users; 25 s for Pro on the web |
| Longer sequences | Extension: +7 s per step, up to about 148 s at 720p | — (Storyboards were for shot planning) |
| Resolution | 720p, 1080p, 4K | 720p (sora-2); up to 1080p (sora-2-pro) |
| Aspect ratios | 16:9, 9:16 | Landscape and portrait |
| Frame rate | 24 fps | — |
| Image to video | Yes | Yes |
| First and last frame | Yes | Not offered |
| References | Up to 3 reference images | Cameos (a recorded likeness) |
| Model tiers | Veo 3.1, Fast, Lite | Sora 2, Sora 2 Pro |
Google describes Veo 3.1's 1080p and 4K output as upscaled, so 4K is best treated as a delivery format rather than proof of extra generated detail.
Audio and dialogue
Sound was the headline for both models. OpenAI launched Sora 2 with synchronized dialogue and sound effects; Google says Veo 3.1 generates sound effects, ambient noise and dialogue natively, and in the Gemini API audio is always on.
Neither vendor published a like-for-like audio comparison, so treat claims that one model's lip-sync is clearly better with caution. What reliably helps with Veo 3.1 is writing the sound into the prompt: the ambience, the key effect, and each line in quotes. An 8-second clip fits a line or two of dialogue.
Realism and physical motion
OpenAI made physical realism Sora 2's signature, with examples like a missed basketball rebounding off the backboard rather than snapping into the hoop. Google's claim for Veo 3.1 is "improved prompt adherence" and more realistic textures than Veo 3. Both are vendor positioning, not a controlled test.
For your own work, the useful test is a scene you care about: physical contact, falling objects, liquid, or two people interacting. Generate it, then watch at full screen for hands passing through objects, sliding feet, or objects changing size.
Clip length and longer stories
This was Sora 2's clearest advantage. Everyone could generate 15-second clips in the app, and Pro users could reach 25 seconds on the web, with Storyboards to plan a clip shot by shot.
Veo 3.1 generates up to 8 seconds at a time. Google's API and Flow can extend a clip in 7-second steps (at 720p), but the simpler route is usually to plan a sequence of short shots and edit them together. If you are coming from Sora, rewrite long prompts as a list of shots, each with one main action and one camera move.
Control over a shot
Here Veo 3.1 gives you more to work with:
- Image to video: start from your own still to lock framing, subject and style.
- First and last frame: add an end image and Veo 3.1 builds the motion between the two.
- Reference images: up to three images of a character, product or place to keep them consistent across separate clips.
Sora 2's equivalent was Cameos, which were about inserting a recorded person into a scene rather than pinning the composition. If your Sora work depended on Cameos of yourself, there is no direct Veo 3.1 equivalent, and Google's safety filters restrict generating real people from photos.
Which is right for you now?
Because Sora 2 is no longer available, the decision is about what to use instead:
- You made short social clips with sound in Sora: Veo 3.1 covers this directly. Use 9:16 for Shorts, Reels and TikTok.
- You relied on 15–25 second single takes: plan them as several 8-second Veo 3.1 shots, or use a model that generates longer clips in one pass.
- You need consistent characters or products across shots: Veo 3.1's start frame and reference images are the better tool.
- You used Cameos of yourself: no like-for-like replacement exists in Veo 3.1.
Trying Veo 3.1 here
Veo Video runs all three Veo 3.1 tiers. Every generation is an 8-second clip with sound, in 16:9 or 9:16, at 720p, 1080p or 4K, from text or from an image with an optional end frame; Veo 3.1 Fast also accepts up to three reference images. Extension and 4- or 6-second lengths are not available here.
New accounts get 30 credits, enough for one Veo 3.1 Lite clip at 720p, so you can test one of your old Sora prompts first. Use Veo 3.1 for final-quality shots. For longer single takes, the AI video generator also runs MiniMax H3 Max.
Related: Veo 3.1 vs Grok Imagine and Veo 3.1 vs Kling AI.
