Moshi
Moshi is an open full-duplex speech-text model that listens and talks at once, with about 200 ms latency, no turn-taking pipeline.
Casts spoken dialogue as speech-to-speech generation on a text LM backbone, with parallel streams for user and system audio; the open answer to GPT-4o voice. Weights released under CC-BY 4.0.
- Date
- Tuesday, 17 September 2024
- Lab
- Kyutai
- Kind
- model
- Access
- open weights
arXiv v1 is 2024-09-17 (the same day as the public release, by memory); Hugging Face repo created 2024-09-11. The ~200 ms figure is from memory.
Sources
This record was checked against its sources on 6 October 2026. How we check