Step 3.5 Flash
196B-total, 11B-active open MoE with 3-way multi-token prediction and 256K context; 74.4% SWE-bench Verified at 100 to 350 tokens/s.
Interleaved 3:1 sliding-window to full attention and MTP-3 cut agent loop cost, and a scalable RL setup spans math, code and tools. Paper figures are 85.4% IMO-AnswerBench, 86.4% LiveCodeBench-v6, 88.2% tau2-Bench and 51.0% Terminal-Bench 2.0. Apache-2.0.
- Date
- Thursday, 29 January 2026
- Lab
- StepFun
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Verified | 74.4% | company |
| Terminal-Bench 2.0 | 51.0% | company |
| tau2-Bench | 88.2% | company |
| Throughput | 100-300 tok/s typical, 350 peak | company |
OpenRouter lists it on 2026-01-29; HF weights from 2026-02-01 UTC; Wikipedia says 2026-02-03; arXiv 2026-02-11. Benchmarks are StepFun-run.
Sources
- arxiv.org/abs/2602.10604
- huggingface.co/stepfun-ai/Step-3.5-Flash
- static.stepfun.com/blog/step-3.5-flash/
This record was checked against its sources on 6 October 2026. How we check