Alibaba officially unveiled Qwen-UI-Agent on August 20, a new GUI agent foundation model spanning mobile, desktop, web, and deep search environments. According to official data, the model achieved 82.1% on the MobileWorld benchmark, outperforming GPT-5.6Sol by 12.0 percentage points and Claude Opus4.8 by 14.6 percentage points. On the AndroidDaily benchmark, it reached a near-perfect 97.5%. On the desktop side, Qwen-UI-Agent scored 79.5% on OSWorld-Verified, surpassing GPT-5.5 and Gemini 3.1 Pro.
文章图片 2
F
文章图片 4
rom an industry perspective, the large model race is rapidly expanding from "chatting and writing code" to "operating interfaces and completing real-world tasks." GUI agents must understand on-screen buttons, input fields, lists, and page navigation, then perform clicks, text entry, and swipes just like a human user. In the past, such capabilities were heavily dependent on simulated environments and often degraded