Grok 4 and Grok 4 Heavy
Grok 4 pairs native tool use with large-scale RL; Grok 4 Heavy runs parallel agents and is first to claim 50% on Humanity's Last Exam.
xAI scaled reinforcement learning on Colossus and added native search and code tools. Heavy reports 50.7% on HLE (with tools) and 61.9% on USAMO 2025; base Grok 4 scores 15.9% on ARC-AGI-2, nearly double Claude Opus 4's 8.6% per xAI.
- Date
- Wednesday, 9 July 2025
- Lab
- xAI
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| Humanity's Last Exam (Heavy, with tools) | 50.7% xAI says first model to reach 50% | company |
| ARC-AGI-2 | 15.9% vs Claude Opus 8.6% per xAI | company |
| USAMO 2025 (Heavy) | 61.9% | company |
| Vending-Bench net worth | $4,694.15 vs Claude Opus $2,077.41 | company |
All headline numbers are xAI-reported. Also in the app for SuperGrok subscribers, with a new SuperGrok Heavy tier for Grok 4 Heavy.
Sources
This record was checked against its sources on 6 October 2026. How we check