WebWorld (8B, 14B, 32B)
Open web-simulator world model (8B, 14B, 32B) trained on 1M+ open-web interactions; training Qwen3-14B on its trajectories lifts WebArena by 9.2%.
A large-scale 'web world model' (8B, 14B and 32B checkpoints) used to train web agents inside simulation. Precursor to Qwen-AgentWorld (2026-06-22), which generalises the idea to many agent environments.
- Date
- Monday, 16 February 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| WebArena gain for Qwen3-14B | +9.2% After training on WebWorld-synthesized trajectories; says comparable to GPT-4o | company |
| Training data / horizon | 1M+ interactions; 30+ steps Authors call it the first open-web simulator trained at scale | company |
Hugging Face repos created 2026-02-13; paper arXiv 2602.14721 submitted 2026-02-16 (10 authors). Claims, including parity with Gemini-3-Pro as a simulator, are the authors' own benchmark results. A follow-up, 'WebWorld: The Browser as a World Model for Self-Improving Web Code' (arXiv 2608.30530), appeared 2026-08-31.
Sources
- api.github.com/orgs/QwenLM/repos?sort=created&direction=desc&per_page=50
- huggingface.co/api/models?author=Qwen&sort=createdAt&direction=-1&limit=60
- arxiv.org/abs/2602.14721
This record was checked and corrected against its sources on 6 October 2026. How we check