Grok-1.5
Grok-1.5 raises context from 8K to 128K tokens and lifts math and code scores sharply over Grok-1.
128K context with perfect needle-in-a-haystack retrieval across the window (xAI test), MATH 23.9% to 50.6% and HumanEval 63.2% to 74.1%. Still behind the best closed models on knowledge benchmarks.
- Date
- Thursday, 28 March 2024
- Lab
- xAI
- Kind
- model
- Access
- app only
Figures
| Measure | Value | Measured by |
|---|---|---|
| MATH | 50.6% vs 23.9% for Grok-1 | company |
| HumanEval | 74.1% vs 63.2% for Grok-1 | company |
| MMLU | 81.3% | company |
| GSM8K | 90% | company |
Announced for early testers on X first; broader rollout to X Premium subscribers followed (exact rollout date not confirmed here; Wikipedia gives 2024-05-15).
Sources
This record was checked against its sources on 6 October 2026. How we check