Grok 4.5
Grok 4.5, trained with Cursor, targets coding and agents. It scores 64.7% on SWE-Bench Pro and 83.3% on Terminal-Bench 2.1, at $2/$6 per million tokens.
A mixture-of-experts coding and agent model built jointly with Cursor on trillions of tokens of Cursor usage data. xAI claims 4.2x fewer output tokens than Claude Opus 4.8 at max effort; a codebase-snapshot leak into training data makes some scores 'optimistic' (reported).
- Date
- Wednesday, 8 July 2026
- Lab
- xAI
- Kind
- model
- Access
- closed API
- Price
- $2 in / $6 out per 1M tokens (under 200K), 2026
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-Bench Pro | 64.7% | company |
| Terminal-Bench 2.1 | 83.3% vs 84.3% for Fable max per xAI | company |
| DeepSWE 1.0 | 62.0% benchmark by Datacurve, run by Artificial Analysis with each provider's harness | third-party |
Sources conflict on the date. DataCamp and Wikipedia say 2026-07-08 (Cursor availability), explainx says Cursor 07-08 and public 07-09, while xAI's post is dated 2026-07-16. Earliest public date used. Cursor reportedly included an earlier codebase snapshot in training data (explainx), so scores are 'optimistic until independent repro'. xAI is branded SpaceXAI since 2026-07-06 per press.
Sources
- x.ai/news/grok-4-5
- www.datacamp.com/blog/grok-4-5
- explainx.ai/blog/grok-4-5-public-launch-spacexai-july-2026
This record was checked against its sources on 6 October 2026. How we check