AI Research Atlas

Chameleon

Meta · 16 May 2024

Early-fusion token-based model that reads and writes interleaved images and text in one transformer (7B and 34B).

Tokenised images into the same stream as text from the start (no separate vision encoder), and reached leading captioning and competitive text results. It is a precursor to native-multimodal training.

Date
Thursday, 16 May 2024
Lab
Meta
Kind
paper
Access
open weights (restricted license)

Figures

MeasureValueMeasured by
Mixed-modal generationmatches or beats Gemini Pro and GPT-4V
human eval on long-form mixed-modal prompts, per authors
company

Release was research-only with image-generation capability withheld from the public checkpoints (not verified from pages opened; arXiv v1 2024-05-16).

Sources

  1. arxiv.org/abs/2405.09818

This record was checked against its sources on 6 October 2026. How we check