PaliGemma
3B vision-language open model pairing a SigLIP image encoder with Gemma 2B; does detection and segmentation as tokens.
Google's first open vision-language model, fine-tunable for captioning, VQA, detection and segmentation (coordinates and masks emitted as special tokens). Announced at I/O 2024 alongside the Gemma 2 preview.
- Date
- Tuesday, 14 May 2024
- Lab
- Google DeepMind
- Kind
- open-weights
- Access
- open weights (restricted license)
Wikipedia lists a different date (2024-07-10); the I/O blog shows the 2024-05-14 announcement.
Sources
- blog.google/technology/ai/google-gemini-update-flash-ai-assistant-io-2024
- developers.googleblog.com/en/gemma-explained-paligemma-architecture/
This record was checked against its sources on 6 October 2026. How we check