Jamba
Jamba is the first production-grade hybrid, with interleaved Transformer and Mamba layers plus MoE, 52B total / 12B active and 256K context on one 80GB GPU.
Shows Mamba blocks can be mixed into a full-scale LLM to cut KV-cache memory and lift long-context throughput while keeping Transformer quality. Established the hybrid template later used by Nemotron-H, Qwen3-Next and Kimi Linear.
- Date
- Thursday, 28 March 2024
- Lab
- AI21 Labs
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parameters (total / active) | 52B / 12B 256K context, fits on one 80GB GPU | company |
arXiv v1 2024-03-28, v2 2024-07-03. Lead authors Opher Lieber, Barak Lenz.
Sources
This record was checked against its sources on 6 October 2026. How we check