AIREITER

Qwen-UI-Agent: Official Benchmarks and How to Access It

Last Updated: 2026-08-20 19:22:38

Qwen-UI-Agent posts the strongest mobile-use results in its reported comparisons - 92.2% on the real-device MobileWorld-Real test, 97.5% on AndroidDaily - but the 27B model behind those numbers has no confirmed public weights and no API. Today you can read the technical report, watch the demos, and download only its smaller MAI-UI predecessors. That gap between leaderboard and availability is the whole story of this release, so this guide rounds up every official score and then audits exactly what you can get your hands on.

The full Qwen-UI-Agent scoreboard

Qwen-UI-Agent is a foundation GUI agent published by Alibaba's Tongyi-MAI team as a technical report on July 30, 2026 (arXiv 2607.28227). One model spans mobile, computer-use, web, and DeepSearch environments, with a unified action space covering GUI operations, CLI execution, and batched action sequences. Across the official results page, Qwen-UI-Agent leads every mobile and every grounding benchmark shown, ranks second on OSWorld-Verified, and ranks third on the OSWorld-v2 long-horizon test - the report evaluates 27B, 35B-A3B, and 4B variants, and the scores below are the 27B flagship's.

Qwen-UI-Agent official benchmark scores compared with Claude Opus 4.8 and Gemini 3.1 Pro

Mobile: the state-of-the-art claim

Mobile is where the report claims state of the art, and the margins are not subtle. On the GUI-only split of MobileWorld, Qwen-UI-Agent's 82.1% sits 8.9 points clear of the next general-purpose model (Seed 2.1 Pro at 73.2%) and 38.2 points clear of the strongest specialized GUI agent (GUI-Owl-1.5-32B-Instruct at 43.9%).

ModelMobileWorld (GUI-only)MobileWorld-RealAndroidDaily
Qwen-UI-Agent (Alibaba)82.192.297.5
Seed 2.1 Pro (ByteDance)73.288.795.2
Claude Opus 4.8 (Anthropic)67.584.793.0
Gemini 3.1 Pro (Google)58.186.293.8
Qwen 3.7 Plus (Alibaba)62.372.779.8

The official tables use GPT-5.6 Sol for the mobile rows (70.1 / 85.4 / 92.6) and GPT-5.5 for the desktop rows - worth knowing before you quote a single "GPT" number.

MobileWorld-Real is the headline benchmark and the authors' own: 409 end-to-end tasks across 104 real Android apps in seven daily-use domains, run on live devices with real accounts and real-world interruptions, scored by a five-VLM majority-vote judge.

Computer use: strong, but not first

Desktop is where the report's "competitive" wording does quiet work. Qwen-UI-Agent's 79.5% on OSWorld-Verified ranks second, 3.9 points behind Claude Opus 4.8. On the long-horizon OSWorld-v2, the 40.0% is a partial-progress score - binary completion is 13.9% - and Claude Opus 4.8 (54.8) and GPT-5.5 (49.5) both sit well above it.

ModelOSWorld-VerifiedOSWorld-v2 (partial)OSWorld-v2 (binary)
Claude Opus 4.883.454.8not reported
Qwen-UI-Agent79.540.013.9
Seed 2.1 Pro78.8not reportednot reported
GPT-5.578.749.513.0
MiniMax M3not reported22.3not reported

The efficiency angle is real, though: batched actions bring the average to 135.8 steps per task, against 326.7 for MiniMax M3 and 173.5 for Qwen 3.7 Plus, while still beating every open-weight model by 17+ points on partial progress. On OSWorld-v2, CLI commands account for more than 50% of the agent's operations - the hybrid GUI+CLI design is doing measurable work here.

Browser and DeepSearch

WebArena goes to Qwen-UI-Agent at 73.6%, ahead of Claude Opus 4.8 (71.9), GPT-5.5 (69.5), and Gemini 3.1 Pro (65.3). On the research side, BrowseComp lands at 64.1 versus 61.0 for the Qwen3.5-27B baseline, and BrowseComp-ZH at 75.0 versus 62.1 - the DeepSearch training transfers, not just the pointing.

GUI grounding: a clean sweep

Grounding - clicking the right pixel from a visual instruction - is a clean sweep: per the official results page, Qwen-UI-Agent ranks first on every grounding benchmark shown.

BenchmarkQwen-UI-AgentBest competitor
ScreenSpot-Pro (no zoom)76.672.9 (GUI-Owl-1.5)
ScreenSpot-Pro (zoom-in)81.580.7 (Seed 2.1 Pro)
SS-V297.596.6 (Seed 2.1 Pro / Qwen 3.7 Plus)
MM-GUI-L292.691.3 (MAI-UI)
OSW-G-R78.578.2 (Qwen 3.7 Plus)
UI-Vision70.068.0 (Qwen 3.7 Plus)

The generalist tax is small

GUI-specialized models usually forget how to be language models. Qwen-UI-Agent mostly doesn't: MMLU-Pro lands at 86.5 against Qwen3.5-27B's 86.0, and IFEval at 90.2 against 90.4. Against the specialized field, the gap is a cliff.

BenchmarkQwen-UI-AgentQwen3.5-27BUI-Venus 30B-A3BGUI-Owl 32BOpenCUA-72B
Tau2-Bench89.989.222.76.114.4
Terminal-Bench 2.050.141.13.20.09.0
BFCL-v474.271.319.832.728.3
BrowseComp64.161.0not reportednot reportednot reported
QwenClawBench44.248.56.45.111.4

The one visible loss to its own baseline is QwenClawBench (44.2 vs 48.5). Note also that the official page marks these broader-capability scores as reproduced in the authors' evaluation environment.

How Qwen-UI-Agent was trained

Per the technical report, Qwen-UI-Agent's training data comes from a real-device fleet - more than 100 physical phones covering 150+ apps - feeding an agent-driven data flywheel that constructs and diagnoses its own tasks. Supervised fine-tuning is followed by multi-stage RL on trajectories exceeding 100 turns, rolled out across more than 10,000 concurrent environments.

How much to trust the scoreboard

Three caveats come straight from the official page, and all three should travel with any number you quote. First, MobileWorld-Real is the authors' own benchmark. Second, dagger-marked baseline scores were reproduced by the authors in their own environment - harness, judge, simulator, and task subsets may differ from the other vendors' official evaluations. Third, the 40.0% OSWorld-v2 figure is partial progress, not success rate. (Source: the official results page notes and the technical report.)

"Alibaba's new GUI agent wins the phone and stalls on the desktop." - @Daily_AI_Brief on X, which also noted that "Alibaba scored every baseline in its own environment."

Independent readings on X reached the same interpretation; a Japanese-language walkthrough by @itarutomy flagged the real-device mobile lead while stressing that the desktop scores are second place. The official announcement thread from Tongyi Lab's own Pengxiang Li (@oliverlee1999) is the primary source for the release itself. Net: treat MobileWorld-Real as a strong claim awaiting reproduction, and treat OSWorld-Verified as the more portable comparison - it is not a benchmark the authors introduced.

How to access Qwen-UI-Agent today

Qwen-UI-Agent is currently a research release: a paper, a project page, and code repositories - but no confirmed public checkpoint of the model itself. Here is what each public artifact contains.

Qwen-UI-Agent official project page with benchmark comparisons

The artifact audit

ArtifactLinkWhat it contains
Technical reportarxiv.org/abs/2607.28227Full paper, submitted July 30, 2026; PDF, HTML, and TeX source
Project pagetongyi-mai.github.io/Qwen-UI-AgentCapability demos, full benchmark charts, citation info
MAI-UI repositorygithub.com/Tongyi-MAI/MAI-UIApache-2.0 repo, retitled for Qwen-UI-Agent; links the MAI-UI-8B and MAI-UI-2B Hugging Face checkpoints released December 2025
Qwen-UI-Agent repositorygithub.com/Tongyi-MAI/Qwen-UI-AgentWebsite source only - the repo states it contains no model, training code, or agent implementation

The MAI-UI repository is the official continuation home: it announced Qwen-UI-Agent on July 30, 2026, and currently links the only downloadable checkpoints in this lineage.

The weights situation

The only downloadable checkpoints are MAI-UI-8B and MAI-UI-2B, the December 2025 predecessors, linked from the MAI-UI repository's Hugging Face references. No Qwen-UI-Agent 27B, 35B-A3B, or 4B checkpoint is confirmed public - availability status here was checked against the official repositories and project page in late August 2026, and none surfaced a release. Do not treat links to MAI-UI checkpoints or the website-source repository as Qwen-UI-Agent model weights.

No API, no self-hosting - yet

No public Qwen-UI-Agent API endpoint, hosted product, or vLLM/Ollama package appears in any official channel. The practical options today: for desktop workflows you must ship, the reported OSWorld results favor Claude Opus 4.8; for open-weight GUI experimentation, MAI-UI 2B/8B plus your own harness is the available route - and it is MAI-UI, not Qwen-UI-Agent. Check the official Tongyi-MAI repositories and Qwen channels for release announcements. For the general-purpose Qwen line that does have API availability, the AIReiter Qwen model catalog tracks current endpoints and pricing.

Where Qwen-UI-Agent sits in the family tree

Qwen-UI-Agent is the third generation of Alibaba's GUI agent line: UI-TARS pioneered the Qwen GUI approach in 2025, MAI-UI (arXiv 2512.22047, December 2025) introduced the real-world-centric data flywheel and open-sourced 2B/8B weights, and Qwen-UI-Agent scales that recipe to 27B with hybrid GUI+CLI execution. Claude Opus 4.8 leads both OSWorld metrics in the tables above.

Qwen-UI-Agent FAQ

Is Qwen-UI-Agent open source?

The code repositories are Apache-2.0 and the report is public, but the model itself is not confirmed open-weight. Only the predecessor MAI-UI-8B and MAI-UI-2B checkpoints are downloadable (Hugging Face, December 2025). No 27B, 35B-A3B, or 4B Qwen-UI-Agent checkpoint had a confirmed public release as of this writing.

Can you run Qwen-UI-Agent locally today?

No verified local setup exists - no checkpoint, no GGUF or Ollama package, no vLLM recipe. Local experiments are limited to MAI-UI 2B/8B with your own harness.

Is there a Qwen-UI-Agent API?

No public endpoint or hosted product has been announced. Check the official Tongyi-MAI GitHub repositories and the Qwen team's channels for release announcements.

How does Qwen-UI-Agent compare with Claude Opus 4.8 for computer use?

Claude Opus 4.8 leads OSWorld-Verified 83.4 to 79.5 and OSWorld-v2 partial progress 54.8 to 40.0, while Qwen-UI-Agent leads WebArena 73.6 to 71.9 and every mobile benchmark shown. Long-horizon desktop work currently favors Claude; mobile-first work is where Qwen-UI-Agent's reported numbers are strongest.

What is MobileWorld-Real?

A real-device Android benchmark built by the Qwen-UI-Agent authors: 409 tasks across 104 live apps in seven daily-use domains, scored by a five-VLM majority-vote judge. Its results are self-reported and await independent reproduction.

What would change the picture

Three signals would materially change this assessment:

SignalWhy it matters
A Qwen-UI-Agent checkpoint under the Tongyi-MAI Hugging Face presenceConverts the research release into a buildable option and triggers third-party numbers
An OSWorld-Verified reproduction outside Alibaba's harnessSettles whether 79.5% holds beyond the authors' environment
Any product surface: API, Qwen app integration, or previewMoves the comparison from paper tables to production trade-offs

Until then, the unresolved trade-off stands: the mobile-use leads in the comparisons published so far are self-reported, on a benchmark the same team built.

Related reading: Alibaba's newer general-purpose flagship and its open-weights status are covered in Qwen3.8-Max open weights, with endpoint costs in the Qwen3.8-Max API pricing guide.