SeedRealtime
Native audio-visual full-duplex LLM that fuses audio, video and text in one stream and can speak up unprompted; rolled out at scale at launch.
One architecture instead of an ASR plus VLM plus TTS cascade. ByteDance's human evaluation says it halves conversational timing problems versus cascaded systems (interruptions, sluggish replies, false triggers).
- Date
- Wednesday, 5 August 2026
- Lab
- ByteDance Seed
- Kind
- model
- Access
- app only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Audio-visual conversational pacing issues vs cascaded systems | reduced by half ByteDance human evaluation | company |
Company-run evaluation; no public benchmark numbers in the post.
Sources
This record was checked against its sources on 6 October 2026. How we check