Mobile UI testing tools in 2026, compared by someone who runs one

Anton Malinski ·

Prices and limits verified 2026-09-29 from each vendor’s public pages; tell us if something changed.

Teams keep asking me which mobile UI testing tool they should use, but the question folds two decisions into one. You choose the framework that drives the app, then you choose where its devices or simulators run. We build Marathon Cloud, so read accordingly; I have tried to keep every product claim tied to what the vendor publishes.

I keep them apart because changing execution infrastructure is usually much less invasive than changing a test framework. An Espresso suite does not become a different suite because it moves from a CI runner to a cloud pool. Most of the expensive mistakes I see start with comparing a framework to a device cloud as if they were substitutes.

Frameworks first

Espresso

Espresso is Android’s native UI testing framework. It works inside the instrumentation stack, understands Android application state, and is the default place I would start for an Android app whose team is comfortable writing tests alongside the product code.

XCUITest

XCUITest is Apple’s native UI automation path through XCTest. It fits naturally with Xcode projects and test plans, and it is the framework I expect when an iOS team owns both the app and its automation.

Maestro

Maestro describes flows in YAML and drives the app from the outside. I reach for it when readable cross-platform flows and quick authoring matter more than keeping the tests in the native application language.

Appium

Appium is the broad, cross-platform option in this group. It is useful when a team wants a WebDriver-style automation interface, already has Appium expertise, or needs one abstraction across different application and platform combinations.

Where AI agents fit

The newest category is a tool you give a goal in plain language, and an LLM taps through the app until it decides the goal is met. Several testing vendors now sell some version of it, and coding agents can drive a local emulator through an MCP server such as Emu. I use this kind of tooling every day, and I would still keep it out of the merge gate.

The reason is the signal. An LLM choosing its next tap is non-deterministic, so the same build can pass on one run and fail on the next for reasons that have nothing to do with the app. A check that runs a thousand times a day has to give the same answer for the same code, or people learn to rerun it until it goes green, and at that point it has stopped checking anything.

Where it does pay off is before the test exists. An agent can explore a flow, reproduce a bug from a vague report, and draft the Espresso or XCUITest code that pins it down. It can also confirm a fix on a device while you work. The output I keep is the deterministic test it helped write, and that test is what runs on every pull request.

The free route first

Gradle Managed Devices

Gradle Managed Devices lets the Android Gradle plugin provision emulators on the build machine, start them from snapshots, restore clean state between tests, cache results, and rerun only tests likely to produce a different result. Device groups run several profiles in parallel, while android.experimental.androidTest.numManagedDeviceShards creates identical virtual instances for each profile; total instances are devices multiplied by shards.

The ceiling is the machine. Google’s docs call sharding resource intensive and warn that provisioning can end in a timeout when Gradle cannot create the requested devices. Automated Test Device images reduce CPU and memory use, but they exist only for API level 30, and screenshot tests that depend on hardware rendering are not supported on those images.

I would also switch off the result caching for UI tests. They talk to mock or real backends that are not in the build graph, so an unchanged input does not guarantee an unchanged outcome, and a cached pass on a broken test is a surprise you find in production. Our Gradle Managed Devices comparison is the longer version of that tradeoff.

Flank

Flank is optional third-party tooling around Firebase Test Lab, not a device cloud of its own. It gives teams a configuration layer for driving FTL, but the execution still lands in Firebase Test Lab, so its practical ceiling remains the capacity and limits of the underlying service.

xcodebuild -parallel-testing

On Apple projects, xcodebuild -parallel-testing can spread tests across workers on Macs you control (see Apple’s Xcode documentation). This is the clean free-software route when the Macs already exist. Its ceiling is still the hosts you own, their simulator capacity, and the maintenance work attached to keeping those hosts ready.

Where the suite actually runs

Firebase Test Lab

Firebase Test Lab runs through the gcloud CLI or Firebase Console; Flank is optional. Android virtual devices are $1 per hour, while physical devices are $5 per hour and iOS runs on physical devices only. Instrumentation runs have a 45-minute limit on physical devices and a 60-minute limit on virtual devices.

Its --num-flaky-test-attempts setting reruns the whole test execution on a device when any test case fails, a fixed number of times. Sharding is exposed through the gcloud beta flags --num-uniform-shards and --test-targets-for-shard. Pick it when you want pay-as-you-go Android virtual devices through a Google Cloud project or need iOS tests on physical devices; compare Firebase Test Lab.

AWS Device Farm

AWS Device Farm runs on real phones and tablets in us-west-2. It starts with 5 concurrent devices, with more available on request, and has a 150-minute hard limit per automated run. Metered use is $0.17 per device-minute ($10.20 per hour), or $250 per month per unmetered device slot.

Pick it when real phones and tablets are the requirement and its self-serve pay-as-you-go or unmetered slot model fits the run pattern; compare AWS Device Farm.

BrowserStack App Automate

BrowserStack App Automate uses real devices only (30,000+ iOS and Android devices). Device Cloud starts at $199 per month per parallel test, billed annually; Device Cloud Pro is $249, with larger volumes handled through sales. The workflow is upload the app to get an app_url, then start the run.

Auto-rerun of failed tests is configured per framework, and sharding is available by setting the shards parameter per run. Pick it when physical hardware such as camera, fingerprint or face sensors, NFC, or OEM-specific behaviour is the point of the test; compare BrowserStack.

Sauce Labs

Sauce Labs has real devices, emulators and simulators, with Appium, Espresso and XCUITest support. Virtual Device Cloud starts at $149 per month and Real Device Cloud at $199 per month, per parallel and billed annually. Its workflow is a saucectl YAML configuration followed by saucectl run, with the app uploaded automatically.

Retries are opt-in through saucectl retries and smartRetry, and saucectl creates sharded jobs once sharding is configured. Pick it when you need Appium or a mix of real and virtual devices under the same service; compare Sauce Labs.

TestMu AI (formerly LambdaTest)

LambdaTest is now TestMu AI. Its automation plans start at $99 per month, while real-device and enterprise tiers go through sales. It offers real devices plus emulators and simulators, and supports Appium, Espresso and XCUITest.

Automatic test splitting is available through HyperExecute. Pick it when you need that framework and device breadth and its plan model fits the team; compare TestMu AI.

Maestro Cloud

Maestro Cloud runs Maestro flows only and costs $250 per device per month. It provides automatic device allocation and parallelization, plus automatic retries for flows suspected to be flaky.

Pick it when the suite is entirely Maestro and you want the cloud built around that flow model; compare Maestro Cloud.

emulator.wtf

emulator.wtf is Android-only and runs Espresso and instrumentation tests. Its sharding choices are targeted-runtime, balanced, even, uniform, or explicit, and --num-flaky-test-attempts supports up to 10 attempts. Outputs include merged JUnit XML, logcat and video, with a web results UI and Gradle plugin.

Pick it when the scope is specifically Android instrumentation and those sharding controls match how you want to divide the suite; compare emulator.wtf.

Xcode Cloud

Xcode Cloud covers iOS, iPadOS, macOS, tvOS, watchOS and visionOS, with no Android. It is available only on Apple-hosted Xcode Cloud and integrates through the App Store Connect API and webhooks. Test-plan repetition can rerun failing tests until they pass.

The included tier is 25 compute hours per month; the paid tiers are 100 hours for $49.99, 250 hours for $99.99, and 1000 hours for $399.99. Pick it when the work is Apple-only and already belongs in Apple’s hosted CI; compare Xcode Cloud.

Marathon Cloud

Marathon Cloud runs Espresso, XCUITest and Maestro on Android emulators and iOS simulators provisioned per run; it is not a physical device farm. It costs $2.00 per hour per Android virtual device and $3.00 per hour per iOS virtual device, billed per second, with 50 free device hours on signup. The pool is sized from the suite’s historical duration to target 15 minutes, which is a target rather than a hard limit, and --concurrency-limit can cap it.

The default per-test timeout is 300 seconds, while total suite duration has no hard cap. A failed test is rescheduled on a different device; a pass on retry is reported as flaky, repeated failure as failed, and recent outcome history drives preventive retries. Runs can start from the official GitHub Action, Bitrise step, Docker image, documented CircleCI integration, any CI, or a laptop.

One CLI command sends the app and test bundle. Results include Allure, JUnit XML, an HTML timeline, per-test screen recording, and per-test device logs. For the documented cost example, a 120-minute suite uses 8 devices for about 15 minutes, or 2 device-hours: $4.00 on Android and $6.00 on iOS. Pick it when native mobile suites on virtual devices, per-test retries, and a pool sized per run are the requirements.

What I’d actually do

I would make four decisions, in this order of importance, and the tool comes third.

1. Choose the checks

Start from what the app must not break: the critical flows, and the production issues you have already had. Collect device usage statistics and crash and bug reports, and turn each real issue into a test. That list, not a vendor’s feature page, decides what the rest of the setup has to support.

2. Decide when each check runs

Every check belongs to one or more stages of the development cycle:

  • local, on uncommitted code, while the engineer is writing it
  • local, as a pre-commit check
  • on CI for every commit
  • on every pull request, before merge
  • on a schedule, every night or every few days
  • before a release

The fast checks run locally on an Android emulator and an iOS simulator on the laptop every SWE and SDET already has; nobody needs a phone on a USB cable to find out that a button moved. The full UI suite runs pre-merge on every pull request, so a regression is caught in the change that caused it. Wider device matrices go on a schedule, and the few checks that need real hardware run before a release.

3. Pick what executes them

This is the product choice, and it follows from the first two. For the pre-merge suite that means many virtual devices in parallel. If you want to start today without choosing a parallel count, the self-serve pay-as-you-go paths are Firebase Test Lab, AWS Device Farm and Marathon Cloud; BrowserStack and Sauce Labs sell per-parallel plans billed annually, TestMu AI starts from a monthly plan, Maestro Cloud asks you to choose a device count, and Xcode Cloud a compute-hour tier.

Start with virtual devices and move to real ones only because of real issues your app has, not because general advice says so. In my experience most of them reproduce on a virtual device once it has the right API level, screen size, locale or configuration. When one does not, that test goes to a real device, in a real-device cloud or a small lab of your own, and that set stays small: camera, biometrics, NFC, manufacturer quirks and whatever else your data says cannot be emulated.

Keep an eye on how device usage changes. The most used device in your analytics today can be obsolete a year from now, so rebuild the device matrix from current statistics instead of carrying last year’s list forward.

4. Pick the frameworks

The framework matters least of the four. Espresso, XCUITest, Maestro and Appium can all express the checks from step 1; pick the one your team writes and maintains well, and confirm the executor from step 3 runs it. A Maestro-only team can use Maestro Cloud or Marathon Cloud, and an Apple-only team already on Xcode Cloud can stay there unless running from another CI is itself a requirement.

Further reading

Our earlier “Android UI tests on every PR” series is the longer, hands-on evaluation behind this post. It predates the 2026-09-29 check, so where prices or limits differ, the numbers above win.