LongCat-Flash-Chat
Meituan's first open LLM: 560B-total MoE that activates 18.6B to 31.3B parameters per token via zero-computation experts; trained on 20T+ tokens in 30 days.
Zero-computation experts spend less compute on easy tokens; shortcut-connected MoE overlaps communication with compute, giving 100+ tokens/s at a stated $0.70 per million output tokens. A multi-stage pipeline targets agentic use. MIT.
- Date
- Saturday, 30 August 2025
- Lab
- Meituan (LongCat)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Total / active parameters | 560B / 18.6-31.3B (avg 27B) | company |
| Training | 20T+ tokens in 30 days | company |
| Inference speed and cost | 100+ tok/s at $0.70 per M output tokens | company |
HF initial commit 2025-08-29 UTC; tech report 2025-09-01; Wikipedia says September 2025. Cost and speed figures are Meituan's.
Sources
- arxiv.org/abs/2509.01322
- huggingface.co/meituan-longcat/LongCat-Flash-Chat
- en.wikipedia.org/wiki/Meituan
- venturebeat.com/ai/chinese-food-delivery-firm-meituans-open-source-ai-model-longcat-flash
- www.ithome.com/0/879/486.htm
This record was checked and corrected against its sources on 6 October 2026. How we check