AI Research Atlas

OmniHuman-1

ByteDance · 3 February 2025

OmniHuman-1 animates a single image of a person from audio (talking, singing, gesturing) with a one-stage, mixed-conditioning diffusion-transformer.

Scales human animation by mixing text, audio and pose conditions during training so weakly-conditioned data is usable. Became the basis for 'talking-head' video in ByteDance's Dreamina and CapCut.

Date
Monday, 3 February 2025
Lab
ByteDance
Kind
model
Access
app only

Only the arXiv page (v1 2025-02-03) was opened; product launch details (Dreamina/CapCut) are from memory and unverified.

Sources

  1. arxiv.org/abs/2502.01061

This record was checked against its sources on 6 October 2026. How we check