AI Research Atlas

Moshi

Kyutai · 17 September 2024

Moshi is an open full-duplex speech-text model that listens and talks at once, with about 200 ms latency, no turn-taking pipeline.

Casts spoken dialogue as speech-to-speech generation on a text LM backbone, with parallel streams for user and system audio; the open answer to GPT-4o voice. Weights released under CC-BY 4.0.

Date
Tuesday, 17 September 2024
Lab
Kyutai
Kind
model
Access
open weights

arXiv v1 is 2024-09-17 (the same day as the public release, by memory); Hugging Face repo created 2024-09-11. The ~200 ms figure is from memory.

Sources

  1. arxiv.org/abs/2410.00037
  2. kyutai.org/

This record was checked against its sources on 6 October 2026. How we check

Related