Ling-3.0-flash
124B-total, 5.1B-active model with a native hybrid-linear design from the start of pretraining, using Kimi Delta Attention and MLA at 5:1; MIT.
Matches or beats Ant's earlier 1T-class Ring-2.6 on key benchmarks at about 12% of its size. Trained with over 10,000 interactive environments; context schedule 8K to 32K to 256K. Reuses Moonshot's Kimi Delta Attention, first published in Kimi Linear (2025-10-30).
- Date
- Sunday, 2 August 2026
- Lab
- Ant Group (inclusionAI)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Total / active parameters | 124B / 5.1B | company |
| TTFT reduction in long-input scenarios | 60% to over 80% with SGLang HiCache and Mooncake | company |
Date from HF weights upload; OpenRouter listed it 2026-07-23. Ling-3.1-flash (560B total, 25B active per OpenRouter) was listed 2026-10-02.
Sources
This record was checked against its sources on 6 October 2026. How we check