Jalapeno inference chip (OpenAI and Broadcom)
OpenAI and Broadcom unveil Jalapeno, an inference-first accelerator taped out in nine months with help from OpenAI models.
Blank-slate design for LLM serving rather than an adapted general accelerator. First results (2026-08-25) on SemiAnalysis InferenceX: about 1.5x higher peak performance per watt and 3.4x lower end-to-end latency than the comparison system on Kimi K2.5, rated at 700 W.
- Date
- Wednesday, 24 June 2026
- Lab
- OpenAI
- Kind
- infra
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| Jalapeno vs comparison system on Kimi K2.5 | about 1.5x performance per watt; 3.4x lower end-to-end latency OpenAI-run on SemiAnalysis InferenceX; comparison system unnamed | company |
| Design-to-tape-out time | 9 months company claim of fastest ASIC cycle | company |
Vendor-run benchmark with an unnamed comparison system; no independent reproduction seen. Part of OpenAI's push to own more of the stack (Cerebras partnership announced 2026-01-14, Broadcom).
Sources
This record was checked against its sources on 6 October 2026. How we check