OmniHuman-1
OmniHuman-1 animates a single image of a person from audio (talking, singing, gesturing) with a one-stage, mixed-conditioning diffusion-transformer.
Scales human animation by mixing text, audio and pose conditions during training so weakly-conditioned data is usable. Became the basis for 'talking-head' video in ByteDance's Dreamina and CapCut.
- Date
- Monday, 3 February 2025
- Lab
- ByteDance
- Kind
- model
- Access
- app only
Only the arXiv page (v1 2025-02-03) was opened; product launch details (Dreamina/CapCut) are from memory and unverified.
Sources
This record was checked against its sources on 6 October 2026. How we check