Logo
FrontierNews.ai

Three Major Physical AI Breakthroughs Show How Robots Are Learning to Think Before They Act

Three leading robotics companies announced major advances in embodied AI that fundamentally change how robots understand and interact with the physical world. Rather than simply predicting what things will look like, these new systems teach robots to understand how physics actually works, enabling them to reason through complex tasks before taking action (Source 1, 2, 3).

What's the Core Problem These New Systems Solve?

For years, embodied AI has suffered from a critical bottleneck: robots could understand language or predict motion, but rarely both at the same time. Traditional vision models trained on video prediction often produce visually convincing results that violate basic physics. A robotic arm might appear to reach correctly on screen but actually intersect with objects, or items might float unnaturally. In the real world, these errors cause failures.

Geek+ addressed this problem by introducing Gravity, a unified embodied AI framework built on what the company calls a "Dual-Brain" architecture. The system combines a "Cognitive Brain" that understands and plans tasks with an "Action Brain" that simulates physical consequences before executing movements. Think of it like how humans work: we think through a problem, mentally simulate the outcome, then act.

Gravity 4D, the framework's foundation model, extends beyond 2D video prediction to learn 3D structure and motion dynamics simultaneously. On a public benchmark called LIBERO-Plus, this approach improved success rates from 73.73% to 78.62% under zero-shot conditions, with the best variant reaching 79.25%. The most significant improvements occurred in scenarios where visual appearance changes but the underlying physics remains constant.

How Are Companies Building the Data Infrastructure for Physical AI?

The second major challenge in embodied AI is data scarcity. Simulations cannot authentically replicate friction, deformation, or unexpected real-world anomalies. Geek+ positioned warehouse picking as the world's largest physical AI training ground, generating hundreds of millions of real picking actions daily across millions of product types. This creates an effectively limitless stream of high-quality training data with unambiguous outcomes: either an item was successfully picked or it wasn't.

ACE Robotics took a different approach with its Ambient Capture Engine 2.0, a human-centric system for capturing high-density physical interaction data. The company introduced the ACE Ego Kit, a lightweight wireless wearable with head, hand, and chest components that includes a specialized glove offering sensitivity of 0.01 newtons. The system synchronizes data from more than 20 heterogeneous sensors with timing error under one millisecond.

ACE Robotics also developed what it calls the "Information-Density Law," classifying embodied data across five levels from L1 to L5. The highest-density L5 data incorporates three-dimensional force and tactile signals, failure-recovery trajectories, and variables from open environments. The company released ACE-Data-0, an L5 household-interaction dataset featuring complex physical tasks and high-precision annotations.

What Practical Capabilities Are These Systems Demonstrating?

Beyond theoretical improvements, these frameworks are enabling real-world robotic capabilities. ACE Robotics' Kairos 3.1 model achieved inference latency of 125 milliseconds on NVIDIA Jetson Thor hardware, enabling near-real-time on-device reasoning. In household laundry scenarios, the system can identify spatial relationships between objects, divide tasks into more than a dozen steps, and verify each stage in real time. If a failure occurs, it identifies the affected step and restarts from that point rather than repeating the entire task.

The model also incorporates self-reflective iteration. When an action fails, the robot evaluates the result and adjusts its strategy. In one test, after an unsuccessful three-finger grasp attempt, the system changed to a four-finger grasp and subsequently completed the task.

LimX Dynamics, which raised nearly $200 million in pre-IPO funding, is developing technology across the complete embodied AI stack, including robotics hardware, model-training infrastructure, and an operating system called COSA. The company organizes its architecture into three integrated layers: System 0 for whole-body motion control, System 1 for vision-language-action models, and System 2 for reasoning and autonomous decision-making. LimX introduced LimX Luna, its full-size interactive humanoid robot, in May 2026, with customer deliveries beginning in China and international markets within one month.

How Are Companies Preparing for Commercial Deployment?

All three companies are moving beyond research prototypes toward industrial-scale deployment. Geek+ announced GINO ECO, an open ecosystem designed to accelerate large-scale commercial deployment of embodied AI. The company positioned itself as an "Embodied AI Solution Expert," creating a systemic loop from model evolution and data accumulation to large-scale deployment.

ACE Robotics introduced three standardized industry solutions designed for deployment within existing commercial workflows. Xiaoman, its integrated fulfillment solution for instant retail, pairs with the new W1 fulfillment robot. The W1 has a robot-to-payload weight ratio of less than 2:1, force-control precision within one newton, and a minimum required aisle width of 75 centimeters, making it suitable for high-density shelving environments. It has been deployed with customers including Sense MartGo, Kuaikeda, and PetroChina convenience stores.

LimX Dynamics plans to use its $200 million funding to deepen integration of high-level cognition with whole-body robotic control, expand manufacturing and delivery capacity, and support international expansion across Europe, the Middle East, and other Asian markets. The company has secured thousands of customer orders since beginning commercial sales, with more than half coming from international markets.

Steps to Understanding the Physical AI Stack

  • Foundation Models: These are large AI systems trained on vast amounts of data that can understand language, vision, and physics simultaneously, serving as the base for specialized robotic applications.
  • Data Infrastructure: High-quality real-world interaction data collected through wearables, sensors, and operational robots creates the training ground for continuous model improvement and generalization.
  • Deployment Systems: Commercial solutions integrate hardware, software, motion control, and task management into standardized packages that can be deployed within days across warehouses, stores, and service environments.
  • Reasoning Layers: Multi-layered architectures separate perception and planning from execution, allowing robots to think through problems before acting, similar to human decision-making processes.

The convergence of these advances suggests that embodied AI is transitioning from a research frontier to an industrial reality.

"In the digital world, a model error may result in a flawed image or paragraph. In the physical world, an incorrect action can have real consequences," stated Wang Xiaogang, Chairman of ACE Robotics. "Kairos 3.1 is built around a first-principles approach to embodied world models, helping robots act more reliably in complex and uncertain environments, and accelerating the arrival of physical AI's Kairos moment."

Wang Xiaogang, Chairman of ACE Robotics

The announcements at the 2026 World Artificial Intelligence Conference also included the launch of PHYSICAL IQ, a unified benchmark for embodied physical intelligence jointly initiated by the Shanghai Artificial Intelligence Association, the Shenzhen Loop Area Institute, and the East China branch of the China Academy of Information and Communications Technology, with participation from more than 20 universities and industry partners.

What distinguishes these 2026 announcements from earlier robotics efforts is the emphasis on unified architectures rather than patchwork solutions. Instead of bolting together separate language models, physics simulators, and motion controllers, companies are building integrated systems where cognition and physical understanding develop together. This represents a fundamental shift in how the industry approaches the challenge of creating robots that can reason about and operate in the real world (Source 1, 2, 3).