AI Research Atlas

Jamba

AI21 Labs · 28 March 2024

Jamba is the first production-grade hybrid, with interleaved Transformer and Mamba layers plus MoE, 52B total / 12B active and 256K context on one 80GB GPU.

Shows Mamba blocks can be mixed into a full-scale LLM to cut KV-cache memory and lift long-context throughput while keeping Transformer quality. Established the hybrid template later used by Nemotron-H, Qwen3-Next and Kimi Linear.

Date
Thursday, 28 March 2024
Lab
AI21 Labs
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
Parameters (total / active)52B / 12B
256K context, fits on one 80GB GPU
company

arXiv v1 2024-03-28, v2 2024-07-03. Lead authors Opher Lieber, Barak Lenz.

Sources

  1. arxiv.org/abs/2403.19887

This record was checked against its sources on 6 October 2026. How we check