AI Research Atlas

Ling-3.0-flash

Ant Group (inclusionAI) · 2 August 2026

124B-total, 5.1B-active model with a native hybrid-linear design from the start of pretraining, using Kimi Delta Attention and MLA at 5:1; MIT.

Matches or beats Ant's earlier 1T-class Ring-2.6 on key benchmarks at about 12% of its size. Trained with over 10,000 interactive environments; context schedule 8K to 32K to 256K. Reuses Moonshot's Kimi Delta Attention, first published in Kimi Linear (2025-10-30).

Date
Sunday, 2 August 2026
Lab
Ant Group (inclusionAI)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Total / active parameters124B / 5.1Bcompany
TTFT reduction in long-input scenarios60% to over 80%
with SGLang HiCache and Mooncake
company

Date from HF weights upload; OpenRouter listed it 2026-07-23. Ling-3.1-flash (560B total, 25B active per OpenRouter) was listed 2026-10-02.

Sources

  1. huggingface.co/inclusionAI/Ling-3.0-flash
  2. openrouter.ai/api/v1/models

This record was checked against its sources on 6 October 2026. How we check

Related