AI Research Atlas

GSPO: Group Sequence Policy Optimization

Alibaba (Qwen) · 24 July 2025

Group Sequence Policy Optimization uses sequence-level importance ratios and clipping to stabilise RL for large MoE models, compared with GRPO.

Defines the importance ratio on whole-sequence likelihood and clips, rewards and optimises per sequence, fixing instability Alibaba saw with GRPO on MoE. The paper says it contributed to the latest Qwen3 models, and sequence-level handling is more tolerant of precision mismatch between training and inference engines.

Date
Thursday, 24 July 2025
Lab
Alibaba (Qwen)
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
PaperarXiv 2507.18071
Submitted 2025-07-24; Qwen blog 2025-07-27
company

Earliest public date is the arXiv submission (2025-07-24). It modifies GRPO from DeepSeekMath. Efficiency and stability claims are the authors' own; adoption in the 2507 Qwen3 releases is stated in the paper's abstract.

Sources

  1. arxiv.org/abs/2507.18071
  2. qwenlm.github.io/blog/gspo/

This record was checked against its sources on 6 October 2026. How we check

Related