ParScale: Parallel Scaling Law for Language Models
Proposes a third scaling axis. Run P parallel, differently transformed passes through one model, which equals about O(log P) more parameters at lower memory growth.
Applies P learnable input transformations, runs forward passes in parallel and aggregates outputs, reusing parameters. The authors report up to 22x less memory increase and 6x less latency increase than parameter scaling for comparable gains.
- Date
- Thursday, 15 May 2025
- Lab
- Alibaba (Qwen)
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parallel scaling efficiency | up to 22x less memory increase Versus parameter scaling at equal quality gain, per the abstract | company |
arXiv 2505.10475 submitted 2025-05-15, 8 authors including Binyuan Hui. Gains are the authors' results on their own pretraining runs; no frontier-scale adoption found.
Sources
This record was checked against its sources on 6 October 2026. How we check