Doubao-1.5-pro
Sparse MoE flagship that activates far fewer parameters than a dense peer. ByteDance says it beats Llama 3.1-405B and gets a 7x performance advantage over dense models.
Training and inference were co-designed. ByteDance reports a MoE matching a dense model with 7x its activated parameters on identical 9T-token data, plus W4A8 serving. Post-training claims no data from other models. Adds native-resolution vision and speech, and was released in the same week as DeepSeek-R1 and Kimi k1.5.
- Date
- Wednesday, 22 January 2025
- Lab
- ByteDance Seed
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| Performance advantage of MoE vs dense (activated params) | 7x at 9T tokens, same data | company |
Source page is the Doubao team's Chinese-language release page dated 2025.01.22. Benchmark tables were images; numbers not captured. Secondary claims of '50x cheaper than GPT-4o' circulated but were not verified.
Sources
This record was checked against its sources on 6 October 2026. How we check