Korean AI Startup Fluiz Achieves Perfect Score on Mobile AI Benchmark, Outpacing Google and Microsoft
A Korean startup called Fluiz has achieved a perfect 100% accuracy score on AndroidWorld, a rigorous mobile AI benchmark developed by Google DeepMind, outperforming AI agents from major tech companies including Google, Microsoft, and Alibaba. The company's mobile AI agent, called FluidGPT, successfully completed all 137 tasks presented in the benchmark, marking the first time any AI system has achieved this level of performance on the platform.
What Makes Mobile AI Agents So Difficult to Build?
Mobile AI agents are programs that run on smartphones and autonomously execute complex tasks on behalf of users, such as operating apps, booking appointments, sending messages, and making financial payments. Unlike desktop environments, mobile screens have diverse structures and frequent transitions as apps open and close, creating a uniquely challenging environment for AI systems to navigate.
The tasks in AndroidWorld go far beyond simply finding and pressing buttons. They require AI systems to recognize screens, understand context, make decisions, and execute long-horizon plans that span dozens of steps across multiple apps while remembering previously viewed information. Because of this complexity, AndroidWorld scores are widely regarded as a true indicator of whether mobile AI agents can move beyond demonstration-level performance and fully complete real-world tasks.
How Does FluidGPT Compare to Competitors?
FluidGPT competed against mobile AI agents developed by several major research organizations and companies. The competition included Artemis, Google's dedicated agent research team, as well as teams from Microsoft Research, Alibaba, and numerous academic institutions and industry players. By solving all 137 AndroidWorld tasks with perfect accuracy, FluidGPT demonstrated capabilities that surpass these well-resourced competitors.
The significance of this achievement lies not just in the perfect score, but in what it represents about the maturity of mobile AI technology. For years, building AI agents that could reliably operate in mobile environments has been considered a top-priority challenge for global tech companies and academia. Fluiz's breakthrough suggests that this challenge is now being solved at a practical level.
Steps to Understanding Mobile AI Agent Capabilities
- Screen Recognition: Mobile AI agents must accurately interpret what appears on a smartphone screen, including identifying buttons, text fields, menus, and other UI elements across thousands of different apps and interfaces.
- Contextual Understanding: The agents need to comprehend the meaning and purpose of what they see on screen, not just recognize visual patterns, so they can make appropriate decisions about what action to take next.
- Multi-App Navigation: Real-world tasks often require moving between multiple apps while maintaining memory of information seen earlier, requiring the AI to track state across app transitions and remember context over extended task sequences.
- Long-Horizon Planning: Complex tasks may require dozens of sequential steps to complete, meaning the AI must plan ahead, anticipate obstacles, and adjust its approach based on unexpected screen changes or error states.
Fluiz was founded by Shin In-sik, a professor at the School of Computing at KAIST (Korea Advanced Institute of Science and Technology). In a statement about the achievement, Shin emphasized the broader implications of the breakthrough.
"Achieving 100% on AndroidWorld proves that Fluiz's FluidGPT technology possesses world-leading accuracy and practicality in the global mobile AI agent market," said Shin In-sik, CEO of Fluiz. "Fluiz will grow into a global AI leader spearheading the on-device AI and mobile operating system ecosystem."
Shin In-sik, CEO of Fluiz
The emphasis on "on-device AI" is particularly noteworthy. This refers to AI systems that run directly on smartphones rather than relying on cloud servers, which could have significant implications for privacy, latency, and user experience. As mobile AI agents become more capable, the ability to run them locally on devices rather than sending data to remote servers becomes increasingly valuable.
Why Does This Matter for the Future of Mobile Technology?
The achievement of perfect accuracy on AndroidWorld represents a watershed moment for mobile AI. For years, AI researchers have struggled with the unpredictability and complexity of real mobile environments. The fact that FluidGPT can now handle all 137 benchmark tasks suggests that mobile AI agents are transitioning from research prototypes to potentially practical tools that could handle real-world smartphone tasks reliably.
This breakthrough could reshape how people interact with their phones. Rather than manually navigating apps and performing repetitive tasks, users could delegate complex multi-step operations to AI agents that understand their intent and execute tasks autonomously. The implications extend beyond consumer convenience to enterprise applications, where mobile AI agents could automate customer service interactions, data entry, and other business processes.
The competition between Fluiz and major tech companies like Google and Microsoft underscores how mobile AI has become a critical frontier in the broader AI race. While large companies have invested heavily in AI research, Fluiz's success demonstrates that specialized startups can achieve breakthrough results by focusing deeply on specific challenges. As the mobile AI market continues to develop, we can expect continued innovation from both established tech giants and emerging startups competing to build the most capable and practical AI agents.