Logo
FrontierNews.ai

The Millimeter Problem: Why Robots Still Can't Grasp the Real World

The gap between what AI understands and what robots can actually do comes down to tiny, invisible errors that pile up in the physical world. While large language models operate in digital space where mistakes are easily corrected, robots face a cascading problem: perception errors, decision-making errors, and motor control errors compound at every stage, and by the time a robot tries to grasp or assemble something, those millimeters of accumulated drift can mean complete task failure.

What's Really Holding Back Physical AI?

At the fifth Global Digital Trade Expo in Hangzhou, Unitree Robotics founder Wang Xingxing delivered a blunt assessment of the robotics industry's core challenge. He explained that the biggest bottleneck in embodied intelligence,the ability for AI to perceive and act in the physical world,is insufficient precision matching between AI model outputs and real-world physics. "The motion looks like it has landed, but that tiny error cannot be auto-corrected, directly causing task failure and a collapse in overall success rate," Wang stated.

This precision problem is fundamentally different from the challenges that generative AI solved. Text-based AI models can tolerate small deviations because information exists in digital space. Robots, by contrast, must sense factory noise and vibration, interpret unpredictable situations, and translate those interpretations into physical actions. Manufacturing, logistics, and shipbuilding sites are emerging as the initial battlegrounds for solving this problem.

When Will Robots Reach Their "ChatGPT Moment"?

Despite the technical hurdles, Wang remains optimistic about the timeline. He predicted that the embodied intelligence sector will reach its own inflection point when robots can complete 80% of tasks in 80% of unfamiliar scenarios through voice-enabled capabilities. At that threshold, he believes global enterprises and national resources will flood into the sector, and this industry tipping point is expected to materialize within the next few years.

The prediction reflects a broader shift in how the world's leading tech companies are approaching physical AI. Google DeepMind is developing "Gemini Robotics" models that enable robots to understand their surroundings and execute actual actions, while collaborating with Boston Dynamics to enhance the AI capabilities of the humanoid robot "Atlas." Nvidia, meanwhile, is building an ecosystem providing the models, simulation tools, and computing infrastructure needed for physical AI development, rather than focusing on robots themselves.

How Researchers Are Accelerating Robot Learning

  • Digital Twin Simulation: By recreating actual work environments in virtual space, robots can learn and verify behaviors under various conditions while reducing costs and safety concerns, since even minor robot malfunctions in real factories can immediately lead to accidents.
  • Human Video Training Data: Researchers are using vast amounts of publicly available human behavior videos from platforms like YouTube as training data, combining this with actual robot behavior data and simulation data to supplement missing physical information like object weight or friction.
  • Cross-Embodiment Learning: The "Open X-Embodiment" project, involving Google DeepMind and other global research institutions, integrated over 1 million real robot trajectory data points collected from 22 types of robots, confirming that experience from one robot type can improve performance across multiple robots with different physical structures.

The Robot Foundation Model (RFM) approach aims to build general-purpose robot intelligence by learning from data collected across various robots and tasks. Instead of training each robot individually for specific tasks, the goal is to extend knowledge gained from multiple robots to new environments and equipment.

National Governments Are Betting Big on Physical AI

The competition for physical AI dominance is no longer confined to private companies. South Korea's government has designated physical AI as one of the "Three Mega Projects for Korea's Great Leap Forward," alongside semiconductors and AI data centers. From 2026 to 2030, the government will invest a total of 1.4131 trillion won (approximately $1.1 billion USD) to develop autonomous factory operation technology in North Jeolla Province and precision control technology for manufacturing processes in South Gyeongsang Province, with demonstrations at actual industrial sites.

"AI security is a systemic engineering challenge, not confined to the model layer but spanning the entire chain of agents, MCP, Skills, and edge devices," said Fan Yuan, chairman of DBAPPSecurity.

Fan Yuan, Chairman at DBAPPSecurity

South Korea's strategy reflects a recognition that having strong AI model technology is only half the battle. The country's next challenge is converting that capability into actual industrial competitiveness. The Ministry of Trade, Industry and Energy is pursuing "M.AX (Manufacturing AI Transformation)," with a goal of building 500 AI factories by 2030.

Corporate partnerships are already expanding across sectors. NC AI is expanding physical AI applications in steel, shipbuilding, and defense sectors through partnerships with POSCO DX, Hanwha Ocean, and Hyundai Rotem, demonstrating how physical AI is moving from research labs into real industrial operations.

The race to solve the millimeter-level precision problem represents a fundamental shift in AI competition. While generative AI captured headlines by mastering language and image understanding, the next frontier is teaching machines to reliably interact with the messy, unpredictable physical world. Whoever solves this problem at scale will likely define the next era of technological leadership.