AI Research Atlas

Veo 3

Google DeepMind · 20 May 2025

Veo 3 generates video with synchronized dialogue, sound effects and ambient audio; Google calls it the first video model with native audio.

First frontier video model to output sound with the picture (speech, effects, music). Launched to US AI Ultra subscribers in the Gemini app and Flow, and on Vertex AI; Gemini API preview 2025-07-17, general availability 2025-09-09 with lower pricing. Triggered the 2025 wave of audio-video models (Sora 2, Kling 2.6, Seedance 1.5).

Date
Tuesday, 20 May 2025
Lab
Google DeepMind
Kind
model
Access
closed API

Figures

MeasureValueMeasured by
SynthID-watermarked items (all Google models)>10 billion
images, video, audio and text, as of I/O 2025
company

"First with native audio" is a company claim. Veo 3.1 followed 2025-10-15; Gemini Omni (2026-05-19) is the successor direction.

Sources

  1. blog.google/technology/ai/generative-media-models-io-2025/
  2. blog.google/technology/ai/google-ai-updates-may-2025/
  3. ai.google.dev/gemini-api/docs/changelog

This record was checked against its sources on 6 October 2026. How we check

Related