Ling-flash-2.0 and Ling-mini-2.0
Ling 2.0 small MoEs with a 1/32 activation ratio. Flash is 100B total and 6.1B active, and Ant claims 7x efficiency over dense models.
First open releases of the Ling 2.0 architecture, guided by Ling scaling laws, with aux-loss-free sigmoid routing, MTP layers, QK-Norm and partial RoPE. Ling-mini-2.0 (16B) appeared on 2025-09-08. MIT.
- Date
- Wednesday, 17 September 2025
- Lab
- Ant Group (inclusionAI)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Total / active parameters (flash) | 100B / 6.1B | company |
| Efficiency vs equivalent dense models | 7x company-claimed | company |
Date from HF upload for flash; mini date from HF repo creation.
Sources
This record was checked against its sources on 6 October 2026. How we check