AI Research Atlas

FACTS Benchmark Suite

Google DeepMind · 9 December 2025

FACTS Benchmark Suite scores LLM factuality across Grounding, Multimodal, Parametric and Search; Gemini 3 Pro leads at 68.8% overall.

Extends the 2024 FACTS Grounding benchmark into four parts, with Search and Parametric error rates dropping 55% and 35% from Gemini 2.5 Pro to 3 Pro. Multimodal scores were the lowest.

Date
Tuesday, 9 December 2025
Lab
Google DeepMind
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
FACTS Score (Gemini 3 Pro)68.8%
overall of four sub-benchmarks; Google-run
company

Built and run by Google with Kaggle.

Sources

  1. deepmind.google/blog/facts-benchmark-suite-systematically-evaluating-the-factuality-of-lar

This record was checked and corrected against its sources on 6 October 2026. How we check