The new Veo Video is live.Contact us
Video with sound, now on Veo 3.1

Veo 3: the Google DeepMind video model that added sound

Veo 3 was the first Veo model to generate sound with the picture: dialogue, sound effects and background noise in the same clip. Google has since replaced it with Veo 3.1, and every Veo video you make here runs on Veo 3.1 — the same kind of prompt, with sound, and closer prompt following.

Try Agent modeNEW

Describe your idea in a conversation and keep refining it.

Try Agent
UploadRequired
Start frame
End frame
  • Images: JPG, PNG or WebP

Production preview

Result

What is Veo 3?

Veo 3 is a video generation model from Google DeepMind, announced at Google I/O in May 2025. Earlier Veo models made silent clips, so any voice or sound had to be added afterwards. Veo 3 generated the audio together with the video: characters could speak lines written in the prompt, and sound effects and background noise arrived with the picture instead of in a separate edit. In October 2025 Google released Veo 3.1, which keeps native audio and improves on Veo 3 in prompt following, sound and photo-to-video quality. Google now lists the original Veo 3 as deprecated and offers Veo 3.1 in its place, in three tiers: Veo 3.1, Veo 3.1 Fast and Veo 3.1 Lite.

Examples

Native audio, now on Veo 3.1

The two clips below are Veo 3.1 clips. They show what Veo 3 made possible — talking characters and scenes with their own sound — as it works today.

Two characters, one scripted exchange

Veo 3.1, image to video. The start frame shows a monkey and a polar bear behind podcast microphones. The prompt scripts their opening: "Welcome back to Bananas & Ice! I am Banana" — "And I'm Ice!" Listen for the two voices and watch the mouths follow each line. This is the trick Veo 3 introduced: write the dialogue in quotes and the model speaks it.

Write a dialogue scene
Veo 3.1 image to video: a monkey and a polar bear hosting a podcast at two microphones

A street at night, sound included

Veo 3.1, made from a reference image. Prompt: "A gentle man is playing the violin by the roadside on a quiet night." The prompt says nothing about sound; listen to the soundtrack Veo builds for the scene on its own. On Veo Video, reference images run on Veo 3.1 Fast, with up to three images per clip.

Veo 3.1 from a reference image: a man playing the violin under a street lamp at night, city lights behind him
Compare

Veo 3 vs Veo 3.1

Veo 3.1 reads the same kind of prompt as Veo 3, so you can paste an old prompt as it is.

Veo 3Google DeepMind, May 2025Runs hereVeo 3.1Google DeepMind, October 2025Generate
ReleasedMay 2025October 2025
Sound generated with the videoYesYes, richer
Follows the promptYesMore closely than Veo 3
Photo to videoYesYes, with better picture and sound
First and last frames—Yes
On Veo VideoReplaced by Veo 3.1Veo 3.1, Veo 3.1 Fast and Veo 3.1 Lite
Prompting

How to write a Veo 3-style prompt with sound

  1. 1

    Set the scene and the camera

    Who is in the shot, where, and what the camera does: "two hosts at a desk, static medium shot" or "handheld, following her through the crowd".

  2. 2

    Write the dialogue in quotes

    Give each speaker a name and a line: Banana: "Welcome back!" Keep lines short enough to fit in 8 seconds.

  3. 3

    Name the sounds you want

    Effects ("a door slams", "rain on the window") and ambience ("café chatter", "distant traffic"). If you leave sound out, the model still generates a soundtrack for the scene.

  4. 4

    Start from a photo if you have one

    With image to video, your photo is the first frame and the prompt only needs the action and the sound.

FAQ

Veo 3 FAQ

Make a video with sound on Veo 3.1

Write a scene and a line of dialogue, or upload a photo to start from.