SAM Audio
Segment-Anything for sound. It isolates a voice, instrument or noise from a mixture using text, visual or time-span prompts (500M-3B params).
First unified separation model prompted by text, visual and temporal spans, runs faster than real time (RTF ~0.7), and ships SAM Audio-Bench plus a reference-free judge model for evaluation.
- Date
- Tuesday, 16 December 2025
- Lab
- Meta
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Inference speed | RTF ~0.7 500M-3B parameters | company |
Sources
This record was checked against its sources on 6 October 2026. How we check