
Alibaba's Tongyi Lab has open-sourced Qwen-AgentWorld, a language world model that acts like a flight simulator for AI agents. Instead of letting an agent crash and burn in a real environment — or even a sandboxed network — this model simulates the environment's next response, letting agents train without the cost and risk of real-world trial and error.
Qwen-AgentWorld covers seven domains, spanning both text and graphical interfaces. For graphical environments like the web, operating systems, and Android, it doesn't bother generating video frames. Instead, it converts observations into code text like HTML and accessibility tree XML, enabling lightning-fast and precise logic simulation. On the AgentWorldBench benchmark, the 397B-parameter mixture-of-experts model (Qwen-AgentWorld-397B-A17B) scored a top average of 58.71, beating GPT-5.4, Claude Opus 4.8, and Gemini 3.1 Pro.
The model offers two key benefits for agent training. First, as a decoupled environment simulator, it can simulate thousands of unseen virtual environments at zero cost — in WideSearch tasks, it matched or even outperformed training with a real search engine. Second, its predictive capability can be internalized as a meta-reasoning mode, so the same model can simulate environmental responses before acting. This led to significant gains in completely new domains: +11.3 on Claw-Eval and +9.0 on the BFCL v4 function-calling benchmark. The model, benchmark, and code are all open-source.