Step 3
321B-total (38B active) vision-language MoE co-designed with its serving system, reaching 4,039 tokens/s per Hopper GPU in decoding vs DeepSeek-V3's 2,324.
Multi-Matrix Factorization Attention shrinks the KV cache; Attention-FFN Disaggregation runs attention and FFN on separate hardware pools. Aimed at cheaper long-context decoding, including on lower-end accelerators. Apache-2.0 weights on 2025-07-28.
- Date
- Friday, 25 July 2025
- Lab
- StepFun
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Decoding throughput per Hopper GPU (4K context, 50ms TPOT, FP8) | 4,039 tokens/s vs DeepSeek-V3 2,324 in the same setup | company |
| Total / active parameters | 321B (VLM) / 38B | company |
Date is arXiv v1 (2025-07-25, during WAIC); HF weights uploaded from 2025-07-28 UTC. Throughput numbers are StepFun's own, theoretical-plus-measured.
Sources
This record was checked against its sources on 6 October 2026. How we check