Logo
FrontierNews.ai

Why AI Still Can't Make a Cup of Coffee: The Embodiment Problem Holding Back AGI

The race toward artificial general intelligence (AGI) hinges on a deceptively simple question: can a machine truly think like a human without a body to interact with the physical world? According to AI researcher Dr. Alan D. Thompson, the answer may be no. While large language models like GPT-4 and Gemini have achieved remarkable cognitive abilities, they remain fundamentally limited by their lack of embodied experience, unable to perform tasks that any average human takes for granted.

What Exactly Is Artificial General Intelligence?

AGI represents a machine capable of understanding and performing at the level of an average human across practically every field, including the ability to interact with the physical world through embodiment. This differs from artificial superintelligence (ASI), which would exceed expert-level human performance in nearly any domain. The distinction matters because it shapes how researchers measure progress toward these milestones.

Thompson defines AGI more strictly than some colleagues: "Artificial general intelligence is a machine capable of understanding the world as well as, or better than, any human, in practically every field, including the ability to interact with the world via physical embodiment." This definition acknowledges that true intelligence requires not just knowledge, but the capacity to act on that knowledge in real-world environments.

Why Can't Today's AI Models Handle Basic Human Tasks?

Consider what the median human can do. According to Thompson's analysis, the typical person in 2024-2025 is a 30-year-old woman from India who speaks two languages, has a working memory of roughly seven items, and can perform practical tasks like making coffee in an unfamiliar kitchen or assembling IKEA furniture. Current AI systems excel at language processing and reasoning but fail at these embodied tasks.

The gap reveals a fundamental truth: intelligence and embodiment are deeply intertwined. All major IQ tests for children under 18 include physical object manipulation, from block-stacking to fine motor skill assessments. These aren't peripheral to intelligence; they're core components of how humans develop and demonstrate cognitive ability.

Thompson points to a thought experiment involving the late physicist Stephen Hawking. While Hawking made groundbreaking discoveries despite severe physical limitations, he had the benefit of decades of full embodiment before his illness. He could manipulate pen and paper, play with ball models, and move his body freely during his formative years. The question becomes: would Hawking have made the same discoveries if he had never experienced embodiment at all?

How Are Researchers Bridging the Embodiment Gap?

Recent advances in robotics offer a glimpse of progress. DeepMind's Gemini Robotics 2, released in July 2026, represents a significant step toward embodied AI. The system demonstrates whole-body humanoid control with advanced multi-finger dexterity and multi-robot coordination capabilities.

  • Dexterity Achievements: On the Apptronik Apollo 2 humanoid with a 22 degree-of-freedom hand, Gemini Robotics 2 can tie knots and seal ziplock bags, with success rates including unscrew bulb at 92 percent, screw bulb at 36 percent, and ziplock sealing at 40 percent.
  • Gripper Precision: On the Franka Duo robot, the system achieved precise insertion tasks at 89.6 percent accuracy and diverse tool kitting at 78.9 percent, demonstrating fine motor control comparable to human hands.
  • Adaptive Learning: The On-Device 2 model can adapt to new robot embodiments with just a few hours of training using fewer than 200 examples, even when robots have drastically different shapes, sensors, and degrees of freedom.
  • Agentic Reasoning: The Embodied Reasoning agent can execute longer task sequences lasting several minutes with hundreds of decisions, self-correcting when steps fail and coordinating multiple robots on workflows a single robot cannot complete alone.

Yet even these advances don't fully solve the embodiment problem. Putting a language model into a robot body doesn't automatically make it smarter; it increases utility and agency. The core intelligence remains limited by the model's training data and reasoning capabilities.

What Does This Mean for the AGI Timeline?

Thompson's assessment places current systems at approximately 98 percent progress toward AGI in terms of robotics capabilities, yet the overall countdown to true AGI remains uncertain. The gap between what AI can do cognitively and what it can accomplish physically continues to narrow, but fundamental challenges persist.

The embodiment requirement isn't merely a technical hurdle; it reflects a deeper truth about intelligence itself. Humans develop understanding through interaction with their environment from infancy onward. We learn physics by dropping objects, spatial reasoning by navigating rooms, and cause-and-effect relationships by manipulating the world around us. An AI system that has never experienced these interactions may struggle to develop the same intuitive understanding, regardless of how sophisticated its language processing becomes.

Steps Toward Embodied AI Development

  • Multi-Sensory Integration: Developing AI systems that can process and learn from multiple sensory inputs including vision, touch, temperature, texture, pressure, and proprioception, mirroring the full range of human sensory experience.
  • Real-World Task Training: Training models on increasingly complex physical tasks in real environments rather than simulations, allowing systems to develop robust understanding of physics, friction, weight distribution, and other embodied knowledge.
  • Adaptive Robot Platforms: Creating flexible robotic systems that can quickly adapt to new embodiments and configurations, enabling AI models to generalize learning across different physical forms and capabilities.
  • Autonomous Decision-Making: Building agentic systems that can plan multi-step sequences, self-correct when failures occur, and coordinate with other agents, moving beyond pre-programmed responses toward genuine autonomous reasoning.

The path to AGI appears to require more than just scaling up language models or adding robotic arms. It demands a fundamental rethinking of how AI systems learn, interact with their environment, and develop the kind of embodied understanding that humans acquire naturally through lived experience. Until AI systems can make that cup of coffee in a strange kitchen, we may still be waiting for true artificial general intelligence.