GSPO: Group Sequence Policy Optimization
Group Sequence Policy Optimization uses sequence-level importance ratios and clipping to stabilise RL for large MoE models, compared with GRPO.
Defines the importance ratio on whole-sequence likelihood and clips, rewards and optimises per sequence, fixing instability Alibaba saw with GRPO on MoE. The paper says it contributed to the latest Qwen3 models, and sequence-level handling is more tolerant of precision mismatch between training and inference engines.
- Date
- Thursday, 24 July 2025
- Lab
- Alibaba (Qwen)
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Paper | arXiv 2507.18071 Submitted 2025-07-24; Qwen blog 2025-07-27 | company |
Earliest public date is the arXiv submission (2025-07-24). It modifies GRPO from DeepSeekMath. Efficiency and stability claims are the authors' own; adoption in the 2507 Qwen3 releases is stated in the paper's abstract.
Sources
This record was checked against its sources on 6 October 2026. How we check