Qwen-AgentWorld (language world models for agents)
Language world models (35B-A3B and 397B-A17B) that simulate agent environments across seven domains, trained on 10M+ interaction trajectories.
Generalises WebWorld to models that predict an environment's next state with long chain-of-thought, trained in three stages (continued pretraining, SFT, RL with rubric-and-rule rewards) and scored on AgentWorldBench. Only the 35B-A3B checkpoint is on Hugging Face.
- Date
- Monday, 22 June 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Training trajectories | 10M+ Seven domains of real-world environment interactions; models 35B-A3B and 397B-A17B | company |
Releasebot dates the blog 2026-06-22; paper arXiv 2606.24597 submitted 2026-06-23 (33 authors). The 'first language world models' claim is the authors'. I did not confirm that the 397B-A17B weights are public.
Sources
- releasebot.io/updates/qwen
- huggingface.co/api/models?author=Qwen&sort=createdAt&direction=-1&limit=60
- arxiv.org/abs/2606.24597
This record was checked against its sources on 6 October 2026. How we check