Gemini 3.1 Pro
Gemini 3.1 Pro scored a verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro, with 80.6% on SWE-bench Verified and 94.3% on GPQA Diamond.
Upgraded core model behind the Deep Think update, shipped in preview for agentic workflows. A separate customtools endpoint favors custom tools over bash. Google's table compares it with Claude Opus 4.6 and GPT-5.2/5.3-Codex, all self-reported.
- Date
- Thursday, 19 February 2026
- Lab
- Google DeepMind
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| ARC-AGI-2 | 77.1% verified; Gemini 3 Pro was 31.1% | ARC Prize |
| SWE-bench Verified | 80.6% single attempt, company-reported | company |
| Terminal-Bench 2.0 | 68.5% Terminus-2 harness, company-reported | company |
| Humanity's Last Exam (no tools) | 44.4% company-reported | company |
Still preview on the API as of 2026-10-04. A Gemini 3.5 Pro was promised for June 2026; the model page still reads "3.5 Pro coming soon" on 2026-10-04.
Sources
- deepmind.google/blog/gemini-3-1-pro-a-smarter-model-for-your-most-complex-tasks/
- deepmind.google/models/gemini/pro/
- ai.google.dev/gemini-api/docs/changelog
This record was checked against its sources on 6 October 2026. How we check