QwQ-32B
32B open-weights reasoning model trained with outcome-reward RL, reported comparable to the 671B DeepSeek-R1 on math, coding and tool-use benchmarks.
Two-stage RL: first math (accuracy verifier) and code (execution server) with outcome rewards, then general capability with reward models and rule-based verifiers. Agentic tool use is folded into the reasoning loop. Apache 2.0.
- Date
- Thursday, 6 March 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
'Comparable to R1' is Alibaba's own benchmark claim (R1 has 671B parameters, 37B active). The blog is dated 2025-03-06.
Sources
This record was checked against its sources on 6 October 2026. How we check