LLaMA
Meta's 7B-65B LLaMA trained on public data only; 13B beats GPT-3 175B on most benchmarks. Weights leaked on 4chan within a week.
Showed that smaller models trained on far more tokens (1-1.4T) can match much larger ones, making strong LLMs runnable on a single GPU. Released to researchers under a non-commercial license on request; the leak (2023-03-03) kicked off the open-weights fine-tune wave (Alpaca, Vicuna).
- Date
- Friday, 24 February 2023
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Sizes / training tokens | 7B, 13B, 33B, 65B; 1.0T-1.4T tokens public datasets only | company |
| LLaMA-13B vs GPT-3 175B | outperforms on most benchmarks paper abstract claim | company |
Official release was research-only by application; the 65B weights leaked via a 4chan torrent on 2023-03-03 (secondary reporting).
Sources
- ai.meta.com/blog/large-language-model-llama-meta-ai/
- arxiv.org/abs/2302.13971
- gigazine.net/gsc_news/en/20230306-llama-65b-leaked
This record was partly confirmed: some claims could not be checked on 6 October 2026. How we check