Llama 3 (8B, 70B)
Llama 3 8B and 70B trained on 15T+ tokens with a 128K-token tokenizer; 400B+ model previewed as still training.
7x more pretraining data than Llama 2 (4x more code), a larger tokenizer that cut tokens by ~15%, and better post-training; Meta AI assistant switched to it across apps. Context stayed at 8K.
- Date
- Thursday, 18 April 2024
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Pretraining tokens | 15T+ 8K context at launch | company |
Sources
This record was partly confirmed: some claims could not be checked on 6 October 2026. How we check