Grok 3 and Grok 3 mini (Think, DeepSearch)
Grok 3 trained on Colossus with about 10x the compute of prior models, adds Think reasoning mode and DeepSearch agent.
xAI's first reasoning model and first agent product. Claims Chatbot Arena Elo 1402, GPQA 75.4% and AIME 2025 93.3% in Think mode (company-run, with xAI's chosen test-time compute). Shipped to X Premium+ first; API followed in April.
- Date
- Monday, 17 February 2025
- Lab
- xAI
- Kind
- model
- Access
- app only
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME 2025 (Think) | 93.3% company chart; test-time compute settings per xAI | company |
| GPQA | 75.4% Grok 3 mini 66.2% | company |
| LiveCodeBench (Think) | 79.4% Grok 3 mini Think 80.4% | company |
| Chatbot Arena Elo | 1402 at launch | company |
Livestream and rollout were 2025-02-17; xAI's blog post is dated 2025-02-19. All numbers are xAI-run; test-time compute settings behind the AIME figure are not stated in the summary I read.
Sources
This record was checked against its sources on 6 October 2026. How we check