arXiv cs.AIPaper
APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
Mobile agents are hard to evaluate because real apps are messy and commercial benchmarks are unreproducible. This trades off both by simulating apps' logic while keeping interactions realistic. Nineteen models tested; none crack 50% autonomous execution yet. This is the benchmark to build on if you're shipping mobile agents, and it signals where the capability gap actually is.