Large Concept Models
Language modelling over sentence embeddings (SONAR, 200+ languages) instead of tokens; up to 7B params on ~2.7T tokens.
Proof-of-concept that predicts the next sentence-level concept in an embedding space, tried with MSE regression, diffusion and quantised variants; reports strong zero-shot multilingual summarisation.
- Date
- Wednesday, 11 December 2024
- Lab
- Meta
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Scale | up to 7B params, ~2.7T tokens authors' proof of concept | company |
arXiv v1 2024-12-11.
Sources
This record was checked against its sources on 6 October 2026. How we check