Qwen Intelligence: What Alibaba Launched, and What Its Own Benchmarks Show
Alibaba's Qwen team launched Qwen Intelligence on 23 September with three "SOTA" mobile agents. Read against its own repositories, the planner's lead is a fifth of a point on a benchmark it built, one agent is July's, and nothing can be downloaded. The one benchmark anyone can run is the most useful part.
Watch
Qwen Intelligence: What Alibaba Launched, and What Its Own Benchmarks Show
~7 min spoken. Keeps playing while you work in another tab.
On 23 September, Alibaba's Qwen team announced Qwen Intelligence, "bringing personal intelligence within everyone's reach", with three agents it describes as state of the art: a Mobile Planner Agent that plans and orchestrates tasks on a phone, a Mobile-Use Agent that carries them out through apps, and a Mobile Creative Agent that turns a sentence into an image. The post also opened a benchmark suite for mobile agents, and linked a website, two GitHub repositories, two project pages, a paper and a leaderboard.
We read all of them. The short version: the headline claims are narrower than the announcement makes them sound, none of the three agents can be downloaded, and the most useful thing released is MobileWorld, a mobile-agent benchmark anyone can run. The planner's own benchmark, despite the announcement's "opening up", can only be taken as a hosted test. The detail follows, agent by agent.
What was launched, and who it is for
qwenintelligence.com is in Chinese only. It offers documentation and a console on Alibaba Cloud, and it sells "full-stack modular delivery" and "device-cloud collaboration": the language of a platform for businesses building phone assistants, not an app for phone owners. When we checked on 29 September, its documentation link returned a "page not found" error. There is no consumer app to install, and nothing in the announcement says when or where there will be one.
That matters for how to read the rest. Qwen Intelligence is, for now, a set of research results and a cloud offering in China. Its agents' claims can be checked only through the reports and benchmarks the team has published.
The planner's "#1": a fifth of a point, on its own benchmark
The announcement says the Mobile Planner Agent is "#1 on MobilePA-Bench, MobilePA-Bench Business & Memory". The figures behind it are in the Qwen-Planner-Agent repository: a 27-billion-parameter planner scores 77.05 overall on MobilePA-Bench, against 76.84 for GPT-6 Astra and 75.71 for Claude Opus 5.
The numbers behind this chart
| Item | Overall score, out of 100 |
|---|---|
| Qwen-Planner-Agent 27B | 77.05 |
| GPT-6 Astra | 76.84 |
| Claude Opus 5 | 75.71 |
| Claude Fable 5 | 74.53 |
| GLM 5.3 | 73.88 |
| Qwen 3.8 Max | 71.77 |
| Kimi K3 | 69.64 |
| Gemini 3.6 Flash | 69.61 |
| Seed 2.1 Pro | 64.89 |
Three things about that lead are worth knowing before repeating it.
It is 0.21 points. MobilePA-Bench has 1,705 tasks, so the gap is the equivalent of three or four tasks. The repository reports no confidence intervals, and a lead of that size would need them before it could be called a ranking at all.
The benchmark is the team's own. MobilePA-Bench was published on 24 August by researchers at the same Alibaba lab, and all eleven of its authors are also authors of the planner's technical report. That is common practice and not a sign of bad faith, but it means the planner was built by the people who defined what the test rewards. Nor can anyone else check it: the benchmark's repository holds only a README, a licence and its website. Its tasks and answers are hidden, and models are scored by submitting a hosted endpoint to the team's private evaluation service.
The benchmark's public leaderboard tells a different story. The MobilePA-Bench leaderboard does not list Qwen-Planner-Agent, and it does not list GPT-6 Astra. For the models it does share with the planner's table, its scores differ, sometimes by several points. Both tables use the same weighting, which we checked against their own figures, so they are two different evaluations of the same models. The repository does not say how its comparison was run. The likeliest reading is that the other models were run inside Qwen's own planning harness, but that is our inference, not something the report states.
| Model | Public leaderboard | Planner's README |
|---|---|---|
| Qwen-Planner-Agent 27B | Not listed | 77.05 |
| GPT-6 Astra | Not listed | 76.84 |
| Claude Opus 5 | 75.52 | 75.71 |
| Claude Fable 5 | 75.31 | 74.53 |
| Kimi K3 | 73.01 | 69.64 |
| Qwen 3.8 Max | 72.51 | 71.77 |
| Gemini 3.6 Flash | 71.21 | 69.61 |
| Seed 2.1 Pro | 63.65 | 64.89 |
Sources: MobilePA-Bench leaderboard; Qwen-Planner-Agent README
Two further details. On the Memory dimension, the planner's own table puts Claude Fable 5 ahead of it, at 76.33 against 74.76, which sits awkwardly beside a claim to be first on "Memory". And we could not find "MobilePA-Bench Business" anywhere public: not in the benchmark paper, not on its leaderboard, not in the planner report's abstract. The comparison also predates the models that shipped in the week of the launch: there is no Claude Fable 5.1, Claude Opus 5.5 or GPT-6 Sol in either table.
The report's other headline may matter more: the planner runs at an estimated $2.41 per 1,000 tasks in output-token cost. The public leaderboard puts Claude Opus 5 at $6.54 and Qwen 3.8 Max at $1.50, but the two are not measured alike - the report's estimate includes thinking tokens and the leaderboard's excludes hidden reasoning - so the figures give an order of magnitude, not a ranking. If a 27-billion-parameter planner is close to frontier models at that price, cost rather than accuracy is its real claim, and a meaningful one for anything running on millions of phones.
The Mobile-Use Agent is July's, and two of its three launch figures are home-grown
The Mobile-Use Agent links to the project page of Qwen-UI-Agent, which the same team announced on 30 July as the successor to its earlier MAI-UI models. The launch gives it a new name; it is not a new model. Its implementation repository, Tongyi-MAI/MAI-UI, reports the scores.
| Benchmark | Score | Built by |
|---|---|---|
| MobileWorld | 82.1% | Qwen's team (Tongyi-MAI) |
| MobileWorld-Real | 92.2% | Qwen's team, "self-built" |
| AndroidDaily | 97.2% in the post, 97.5% in the README | Not stated |
| OSWorld-Verified | 79.5% | Independent |
| WebArena | 73.6% | Independent |
| ScreenSpot-Pro | 81.5% | Independent |
Sources: MAI-UI README; Qwen on X
Two of the three numbers in the announcement come from benchmarks the team built itself: MobileWorld, 201 tasks across 20 apps, and MobileWorld-Real, which the repository calls "a self-built real-device benchmark". The third, AndroidDaily, is given as 97.2 in the announcement and as 97.5 in the repository. The difference is small; that the two disagree at all is the kind of thing a launch post should not get wrong.
The more persuasive evidence is the set of scores on benchmarks Qwen did not build: 79.5% on OSWorld-Verified, 73.6% on WebArena and 81.5% on ScreenSpot-Pro. The repository describes these as competitive with or better than Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.6 Sol. Those are credible numbers for a GUI agent. They were also published two months ago, against models that have since been replaced.
The "90% end-to-end success rate" in the announcement is not tied to any benchmark we could find.
The Creative Agent's link is a paper about something else
The announcement says the Mobile Creative Agent produces an image in three seconds, "about 2x faster than leading peers", and links to an August paper on training pixel-space image models. The paper is a careful study: it finds a recipe that lets pixel-space models match latent-space ones while running 3.18 to 4.75 times faster end to end. But its comparison is between its own models trained two ways. It does not describe an agent, and it does not measure anyone else's image generator. Nothing we found supports "2x faster than leading peers".
What you can actually get
| Item | What is public | Weights or code |
|---|---|---|
| Mobile Planner Agent | Report and website | No |
| Mobile-Use Agent | Website; July predecessor's repository | Predecessors only (MAI-UI 8B and 2B) |
| Mobile Creative Agent | An image-model paper | No |
| MobileWorld | Full benchmark, Apache 2.0 | Yes |
| MobilePA-Bench | Website; tests run on Alibaba's server | No |
| MobileWorld-Real and -Safety | Named in the post | Not found |
Sources: Tongyi-MAI on GitHub; Tongyi-MAI on Hugging Face
Neither agent repository contains a model. Both say so at the top: the planner's repository "is not the release repository for model weights, training code, or the agent implementation", and Qwen-UI-Agent's is "website source only". None of the three agents has released weights. The nearest are the 8-billion and 2-billion-parameter MAI-UI models the team put on Hugging Face last December, the Mobile-Use Agent's predecessors.
Of the four benchmarks the announcement says it is opening, one can be run by anyone today. MobileWorld is on GitHub under Apache 2.0 with its code, Docker environments and a trajectory viewer that lets you inspect each model's run step by step. MobilePA-Bench carries the same licence, but its repository is a website: the tests stay on Alibaba's server. We found no public release of MobileWorld-Real or MobileWorld-Safety.
Who should care
If you build agents for phones, MobileWorld is worth your time now: long, cross-app tasks in emulators you can run yourself, with a published leaderboard to compare against. MobilePA-Bench tests something MobileWorld does not - tool use, memory, reusable skills and handing work to sub-agents, in an environment whose state changes as the agent acts - but you can only reach it by submitting a hosted agent to Alibaba's evaluation service.
If you are choosing a model for a mobile assistant, Qwen Intelligence is not yet something you can choose outside Alibaba Cloud in China. The useful signal is the public MobilePA-Bench leaderboard, where Claude Opus 5 and Claude Fable 5 lead and Kimi K3 and Qwen 3.8 Max are within about three points at lower cost.
If you track Alibaba, the pattern is worth noting: a coordinated launch that packages a two-month-old agent, an unreleased planner and an image-model paper as one product, with its own benchmarks as the evidence. The research underneath is substantial. The launch claims outrun it.
What would change this
Three things would make Qwen Intelligence easier to take at its word: the planner added to the public MobilePA-Bench leaderboard it tops in its own report, with the method of the report's comparison stated; the newest frontier models run on the same harness; and anything from the launch released in a form other people can test. MobileWorld shows the team knows how to do the last of these; the launch did not.