Veo 3
Veo 3 generates video with synchronized dialogue, sound effects and ambient audio; Google calls it the first video model with native audio.
First frontier video model to output sound with the picture (speech, effects, music). Launched to US AI Ultra subscribers in the Gemini app and Flow, and on Vertex AI; Gemini API preview 2025-07-17, general availability 2025-09-09 with lower pricing. Triggered the 2025 wave of audio-video models (Sora 2, Kling 2.6, Seedance 1.5).
- Date
- Tuesday, 20 May 2025
- Lab
- Google DeepMind
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| SynthID-watermarked items (all Google models) | >10 billion images, video, audio and text, as of I/O 2025 | company |
"First with native audio" is a company claim. Veo 3.1 followed 2025-10-15; Gemini Omni (2026-05-19) is the successor direction.
Sources
- blog.google/technology/ai/generative-media-models-io-2025/
- blog.google/technology/ai/google-ai-updates-may-2025/
- ai.google.dev/gemini-api/docs/changelog
This record was checked against its sources on 6 October 2026. How we check