Chameleon
Early-fusion token-based model that reads and writes interleaved images and text in one transformer (7B and 34B).
Tokenised images into the same stream as text from the start (no separate vision encoder), and reached leading captioning and competitive text results. It is a precursor to native-multimodal training.
- Date
- Thursday, 16 May 2024
- Lab
- Meta
- Kind
- paper
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Mixed-modal generation | matches or beats Gemini Pro and GPT-4V human eval on long-form mixed-modal prompts, per authors | company |
Release was research-only with image-generation capability withheld from the public checkpoints (not verified from pages opened; arXiv v1 2024-05-16).
Sources
This record was checked against its sources on 6 October 2026. How we check