Transformers are SSMs (Mamba-2)
Proves attention and state space models are two views of structured semiseparable matrices; Mamba-2's SSD layer runs 2-8x faster than Mamba.
The state space duality (SSD) framework lets SSMs use matrix-multiply hardware and lets linear-attention tricks transfer both ways. Mamba-2 became the building block of most hybrids (Jamba-style, Nemotron-H, Zamba) and the base for Gated DeltaNet and Mamba-3.
- Date
- Friday, 31 May 2024
- Lab
- Princeton / Carnegie Mellon
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Speed vs Mamba | 2-8x faster refined selective SSM | authors |
Tri Dao and Albert Gu; ICML 2024; arXiv v1 2024-05-31. Code and checkpoints were released with the paper (not re-verified on this pass).
Sources
This record was checked against its sources on 6 October 2026. How we check