AI Research Atlas

Insights into DeepSeek-V3: hardware co-design paper

DeepSeek · 14 May 2025

DeepSeek lays out the hardware limits it hit training V3 on 2,048 H800s and argues for low-precision units and better interconnects.

Reflection paper on MLA, MoE, FP8 training and a multi-plane network topology, framing V3's design as hardware-aware co-design under export-controlled GPUs, with proposals for future accelerators.

Date
Wednesday, 14 May 2025
Lab
DeepSeek
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
Training cluster2,048 NVIDIA H800 GPUs
V3 training
company

Sources

  1. arxiv.org/abs/2505.09343

This record was checked against its sources on 6 October 2026. How we check

Related