LongCat-Flash-Thinking
560B reasoning model trained with an asynchronous RL system (DORA) and a domain-parallel recipe covering math, formal proofs and agentic tool use.
Training has two phases, a long-CoT cold start and then large-scale RL. Adds formal and agentic reasoning (theorem proving, tool use). MIT.
- Date
- Sunday, 21 September 2025
- Lab
- Meituan (LongCat)
- Kind
- open-weights
- Access
- open weights
Date from HF commit history; no benchmark numbers recorded.
Sources
This record was checked against its sources on 6 October 2026. How we check