MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World TasksExplained for Beginners
Yi Zhu, Xiongwei Wu, Qiyi Wang +8 more
Abstract
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning 13 functional domains and 212 realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: (1)~Sub-agent Collaboration---decomposing a complex task and delegating specialized work to capable sub-agents; (2)~Memory Usage---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and (3)~Skill Usage---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.
The Problem: Why Another Benchmark Matters
If you’ve followed the rapid rise of AI assistants on phones, you’ve probably noticed a curious gap between demo hype and everyday reality. One day your phone’s LLM can draft a witty text message; the next, it can’t seem to find your calendar, asks for permission to read contacts it already has, or gets stuck in an endless loop when a simple task like “book a dinner reservation” requires three different apps to talk to each other.
Researchers who study these “mobile planner agents” face the same frustration. To make agents that actually work in the real world, they need rigorous ways to test them. But the field’s existing evaluation tools fall into two disappointing camps, each with a blind spot nearly as big as the problem they’re trying to solve.
The first camp consists of GUI-centric benchmarks. These show a model a screenshot of a phone screen and ask it to tap icons or type into fields. It’s like judging a car’s drivability by how shiny its paint is: the model gets good at locating pixels, but it doesn’t learn how to handle the actual engine, transmission, or traffic laws. GUI benchmarks overlook what makes a smart agent truly useful: the ability to run background tools, coordinate across apps, remember what the user liked last week, and recover gracefully when the operating system says “access denied.” A model that can perfectly locate a button on screen may completely fail when it needs to invoke a backend API that doesn’t have a visual element, or when a permission dialog pops up and breaks its flow.
The second camp is static function‑calling benchmarks. These present the model with a list of available API calls and ask it to output the correct JSON string. The model sits in a quiet room, matches a parameter to a function name, and gets a gold star if the format is right. It’s a neat test of formatting ability, but it’s entirely offline. There’s no operating system, no permission prompts, no database that changes after each call, and no unexpected error that forces the model to rethink its plan. In the real world, an agent’s second call might fail because the first call already modified the state, or because the user revoked a permission mid-task. Static benchmarks simply don’t see these dynamics.
MobilePA-Bench was built precisely to bridge this chasm. Its creators noticed that as on‑device LLMs evolve into personal copilots, the mobile OS has become a key testbed—and yet we were evaluating these agents as if they lived in a vacuum. The paper’s abstract puts it bluntly: current benchmarks either test surface‑level screen manipulation or rely on offline API matching detached from real runtime constraints. Neither captures the messy, stateful, permission‑driven reality of using a smartphone as a personal assistant. By bringing a live, interactive sandbox into the evaluation loop, MobilePA-Bench forces models to confront the same frictions real users face, and in doing so, it reveals exactly where today’s best models still stumble.
How It Works: The Technical Mechanics, Explained with a Project‑Manager Analogy
Imagine you’re a project manager (the “central planner”) tasked with organizing a product launch. You have a list of contractors (sub‑agents), a folder of client preferences (the user’s memory), and a set of standardized workflow templates (pre‑packaged skills). Your job is to look at the high‑level goal, decide which contractor should handle which piece, remember the client’s favorite colors from past projects, and pull out a reusable “boilerplate” workflow instead of designing every single step from scratch.
Now, suppose the project site has rules: certain contractors can only enter after a permit is approved, some tasks must be done in a specific order, and if you miss a step, the whole site can grind to a halt. That’s the environment your planner operates in.
MobilePA‑Bench formalizes exactly this kind of scenario, but for a smartphone. At its core is an interactive, stateful sandbox—a simulated mobile operating system that actually executes the actions the model proposes, then reports back what happened. The sandbox isn’t a static checklist; it’s a living environment with a persistent database of contacts, calendar entries, app states, and permission flags. Every time the planner “clicks” a tool, the sandbox updates the database, logs the result, and sends structured feedback back to the model. This feedback includes three pieces: a status (success or failure), an error type (e.g., PermissionDenied, MissingParameter), and a payload (the new state of the database, or a helpful error message).
The action space the planner can choose from falls into four categories, each targeting one of the benchmark’s core capabilities:
-
Basic Tool Use – Direct API calls that query or mutate the phone’s state. Think of these as the everyday utilities: sending a text, checking the weather, adding a contact, adjusting volume. The model must ground the natural‑language request into the correct function name and arguments, respect the order in which tools must be called (e.g., you usually need to request location permission before you can query the GPS), and recover if the sandbox replies “permission denied.” The benchmark injects realistic obstacles: a missing phone number in the request, a permission block that wasn’t anticipated, a database conflict that only appears after two steps. The model has to read the feedback, realize something went wrong, and revise its plan on the fly.
-
Sub‑agent Collaboration – Some tasks aren’t solvable by a single API call. Maybe the model needs to extract text from an image, fill out a form that has no API, or compare two versions of a contract displayed on screen. In MobilePA‑Bench, the central planner can hand off such work to specialized sub‑agents. One sub‑agent might specialize in GUI navigation (tapping, scrolling, reading text), another in visual question answering, another in image processing. The planner’s job is to decide when a task has crossed the threshold into “needs a human‑like visual interpreter” and to route it correctly, packaging along the handoff the necessary context (e.g., “find the ‘Save’ button in the top‑right corner of the consent form”). Success isn’t measured by whether the sub‑agent perfect‑ly completes the task, but whether the planner made the right routing decision and gave clear, complete instructions.
-
Memory Usage – Real users don’t start from zero every time they talk to their assistant. “Remind me to pay the same electric bill I paid last month” or “Send the photo I shared with Mom last year” implicitly reference a personal history. MobilePA‑Bench tasks in this dimension are built from coherent user‑profile worlds that capture long‑term habits, preferences, and past actions. The model must actively query a “search user memory” tool, extract the relevant ID or preference, and weave that information into subsequent planning steps. If the model forgets to check memory, or retrieves the wrong entry, the task fails. The benchmark evaluates this as a gate: all required gold memory IDs must be returned and incorporated for the task to pass.
-
Skill Usage – Planning every single step from scratch is expensive and error‑prone, especially for multi‑step workflows like “book a flight and hotel for a long weekend, then email the itinerary to my team.” Instead of generating a dozen individual API calls, the model can invoke a pre‑packaged skill—a reusable composite procedure that bundles many steps under one function name. The skill loader expands the action space dynamically, exposing the concrete tools the skill relies on, and the model then follows the skill’s downstream instructions. The benchmark tests whether the model knows when a skill is appropriate, loads the correct one, and successfully executes the bundled workflow. It also tests “mixed tool‑skill routing,” where the model might start with a skill, hit a snag, and switch to individual tools to finish the job.
What makes the sandbox particularly powerful for research is the controlled environmental friction baked into every task. Before a model even starts, the sandbox can inject obstacles: a required permission that’s been blocked, a contact name that’s slightly misspelled, a database entry that creates a duplicate if not handled carefully. These aren’t artificial trick questions; they’re realistic frictions that reflect how actual mobile systems behave. Because the sandbox updates a shared backend database in real time, every action has a lasting effect, and the model must track those effects across dozens of steps. This setup also makes the benchmark highly suitable for reinforcement learning: researchers can replay successful and failed trajectories, use them as training signals, and watch how models improve their error‑recovery strategies over time.
Key Results & Benchmarks: What the Numbers Actually Mean
MobilePA‑Bench doesn’t just exist as a theoretical framework; it was put through its paces with some of the most capable large language models available today. The results make for sobering reading, and more importantly, they translate into plain‑language impact that anyone who’s used an AI assistant can relate to.
The benchmark encompasses 1,705 real‑world tasks spread across 13 functional domains and 212 realistic tools. Domains range from “Audio & Entertainment” (media control, camera, screen recording) to “Security & Privacy” (screen lock, biometrics, password vault), with smaller buckets for “Time Management,” “Network & Connectivity,” “Travel & Lifestyle,” and others. Each task is annotated with its initial sandbox state, the set of candidate actions available at the start, and gold‑standard targets for which capability dimension it tests.
Across all four capability dimensions, the overall benchmark score is a weighted aggregate: 50 % Basic Tool Use, 10 % Sub‑agent Collaboration, 20 % Memory Usage, and 20 % Skill Usage. The weighting reflects the authors’ view that stateful, reliable tool execution is the foundation—without it, the more advanced reasoning dimensions can’t even get off the ground.
When the authors ran their evaluations, the headline number was that even the best‑performing frontier LLM reached only 75.52 % overall accuracy. On the surface, that sounds fairly good—three‑quarters of tasks solved. But translate that into the benchmark’s terms, and it means that out of 1,705 tasks, the model got roughly 420 wrong. And the failures weren’t random, evenly‑distributed mistakes; they clustered around the very frictions the benchmark was designed to surface.
Let’s break down how the model fared across the four dimensions:
-
Basic Tool Use (50 % weight): This is the bedrock. The model’s performance here set the tone for everything else. Even on these foundational tasks—selecting the right API, grounding arguments, handling permission blocks, and recovering from errors—accuracy dropped significantly when the sandbox imposed strict tool ordering or runtime permission limits. A typical failure mode: the model tries to read a contact’s phone number without first requesting the “read contacts” permission, the sandbox returns
PermissionDenied, and the model either retries endlessly or abandons the task. In plain language, this is like an assistant that insists on accessing your location every time you ask for the weather, even after you’ve denied it twice. -
Sub‑agent Collaboration (10 % weight): Routing tasks to the right specialist proved challenging. The model often either tried to solve GUI‑heavy tasks using pure API calls (ignoring the need for visual ground‑truth) or routed to the wrong sub‑agent, sending unclear handoff payloads. Success here meant the planner correctly identified when a task required visual form‑filling, dispatched a GUI sub‑agent with a complete instruction set, and then resumed its own planning based on the sub‑agent’s feedback. The 75.52 % overall score already factors this in, but the sub‑agent dimension alone showed wider variance across models.
-
Memory Usage (20 % weight): Tasks that required recalling user profiles, past preferences, or historical habits exposed the model’s short‑term memory limitations. When a request was phrased implicitly—e.g., “Book the hotel I usually pick for conferences in New York”—the model had to query the memory tool, extract the correct preference ID, and use it to inform the booking API. Many models either skipped the memory query entirely or retrieved a generic default, leading to incorrect bookings. The benchmark’s memory gate is strict: all required gold memory IDs must be present in the trajectory for the task to count as successful. This dimension highlighted that even state‑of‑the‑art models struggle to consistently incorporate persistent personal context.
-
Skill Usage (20 % weight): Invoking pre‑packaged composite skills was another area where the model’s planning horizon worked against it. The benchmark includes skills like “synchronized travel and accommodation scheduling,” which bundle many API calls under one function. Models that understood when to use a skill versus planning step‑by‑step performed better, but many defaulted to granular tool calls, accumulating errors as they went. The joint success metric here is evaluated in two modes: Skill‑Only Routing (SOR), where the model must load the gold skill and complete the downstream task using only the skills’ internal tools, and Mixed Tool‑Skill Routing (MTSR), where the model can switch between skills and individual tools as needed. Even the best models missed the optimal routing in a significant fraction of these tasks.
Perhaps the most telling insight from the results is how sharply performance drops when the sandbox introduces strict tool ordering, permission limits, and unexpected runtime errors. In a clean, offline function‑calling test, the same models might score 90 %+ because there’s no state to track, no permissions to manage, and no feedback loop. But put them in MobilePA‑Bench’s interactive sandbox, and the score gap narrows dramatically. The authors emphasize that this drop isn’t a reflection of the models’ language understanding—these are after all LLMs with impressive conversational abilities—but rather their inability to maintain coherent state across a dynamic environment, reason about permission boundaries, and adapt plans when the system pushes back.
The paper’s authors also open‑sourced their complete infrastructure: all 1,705 benchmark tasks, the evaluation datasets, and the high‑throughput sandbox environment. This means other researchers can immediately start running their own models against MobilePA‑Bench, generate failure analyses, and contribute to a growing pool of knowledge about how today’s agents handle real‑world mobile frictions. It’s a rare moment in benchmark publishing where the tooling itself is as valuable as the numbers.
Why It Matters: Key Takeaways
MobilePA‑Bench isn’t just another leaderboard to watch; it points toward several concrete directions that will shape how mobile AI assistants evolve in the near term. Here are the most important takeaways:
-
Real‑world frictions are the bottleneck, not language ability. The steep performance drop when models enter the interactive sandbox tells us that the limiting factor for useful mobile agents isn’t their ability to generate plausible‑sounding API calls or GUI taps. It’s their capacity to track state, respect permissions, and recover from errors in real time. Any framework for building better mobile assistants needs to prioritize state‑tracking modules, permission‑aware policy engines, and error‑recovery loops even more than prompt‑engineering tricks.
-
The foundation matters most. The 50 % weighting for Basic Tool Use in the overall score is deliberate. If a model can’t reliably execute individual tools under realistic constraints, the more advanced capabilities—memory, skill routing, sub‑agent handoffs—become moot. Investing in robust tool‑use infrastructure (clean function schemas, deterministic execution environments, clear feedback signals) will yield bigger gains than chasing ever‑larger model parameters.
-
Memory and skills are the next frontier for personalization. The fact that even top models struggle with the Memory and Skill dimensions means there’s a huge opportunity for improvement. Future assistants that can seamlessly recall “the hotel I stayed at in San Francisco last October” or invoke a “book‑my‑trip‑and‑email‑the‑itinerary” skill will feel dramatically more helpful. The benchmark’s open infrastructure makes it easy to experiment with hybrid approaches: e.g., using retrieval‑augmented generation to augment the model’s memory, or learning when to switch from a skill to fine‑grained tools based on intermediate feedback.
-
Open‑source sandboxes accelerate collective progress. By releasing the full MobilePA‑Bench suite—tasks, sandbox code, evaluation scripts—the authors are lowering the barrier for anyone who wants to test their own agent, conduct ablation studies, or contribute reinforcement‑learning rollouts. This kind of shared infrastructure is rare in AI benchmarking and could become a de facto standard for mobile agent evaluation, much as WebArena or SWE‑Bench have done for web and software‑engineering tasks.
What to watch next: Keep an eye on how the research community iterates on MobilePA‑Bench. We’ll likely see versions that add multi‑modal inputs (voice, camera), tighter integration with real phone ecosystems via Android’s APIs, and new capability dimensions such as “continuous learning” (does the agent update its internal user model as it interacts?). The current results already suggest that the gap between demo‑room fluency and everyday usefulness is narrower than we thought—but also that there’s a lot of low‑hanging fruit in state management, permission handling, and skill composition. If you’re building or evaluating mobile AI agents, MobilePA‑Bench is quickly becoming the place to look for the honest, granular feedback that actually matters.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →