SWE-2
SWE-2, post-trained from the 2.8T-parameter Kimi K3 with multi-trillion-parameter RL, scores 50.0% on FrontierCode 1.1 Main.
Cognition says RL was scaled to multi-trillion parameters for the first time, using Pareto-informed cost penalties so all effort levels train in one run. It claims parity with Fable 5.1 at 64% lower cost; ships in Devin Desktop, CLI, Web and Fusion.
- Date
- Thursday, 10 September 2026
- Lab
- Cognition
- Kind
- model
- Access
- app only
Figures
| Measure | Value | Measured by |
|---|---|---|
| FrontierCode 1.1 Main | 50.0% within one point of Fable 5.1 at 64% lower cost (company) | company |
| Terminal-Bench 2.1 | 92.8% DeepSWE 1.1: 73.0% | company |
Vendor-run evaluations on a vendor-built benchmark. Model, base and benchmarks as stated by Cognition only.
Sources
This record was checked against its sources on 6 October 2026. How we check