MusicLM
Google Research paper generating minutes-long 24 kHz music from text prompts, plus the 5,500-pair MusicCaps dataset.
Treats text-to-music as hierarchical sequence-to-sequence modelling, producing coherent audio over several minutes and conditioning on a hummed or whistled melody plus a style caption. Outperformed earlier systems on quality and text fidelity in the authors' evaluation.
- Date
- Thursday, 26 January 2023
- Lab
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Output sample rate | 24 kHz coherent over multiple minutes | company |
| MusicCaps dataset | 5,500 music-text pairs expert-written captions, released publicly | company |
Paper and samples only at release; the Lyria model line (2023-11) is the DeepMind successor.
Sources
This record was checked against its sources on 6 October 2026. How we check