Why Today's Vision Language Models Still Fall Short of Human Intelligence
Vision language models (VLMs) like GPT-4V and Google's Gemini can see and analyze images, but they remain fundamentally limited compared to human intelligence because they lack physical embodiment and multi-sensory experience. According to AI researcher Dr. Alan D. Thompson, current multimodal systems lack the embodied experience and sensory grounding that humans develop over a lifetime, making them unable to truly understand the physical world the way we do.
What Critical Capabilities Are Vision Language Models Missing?
Vision language models represent a significant leap forward in artificial intelligence (AI). These systems combine language processing with image recognition, allowing them to describe what they see, answer questions about images, and reason about visual content. However, experts argue that current VLMs are missing fundamental capabilities that humans take for granted.
According to Thompson's analysis, the gap between current AI systems and true human-level intelligence involves more than just raw processing power or training data. The missing pieces include sensory and physical experiences that shape human cognition from birth. Thompson defines artificial general intelligence (AGI) as "a machine capable of understanding the world as well as, or better than, any human, in practically every field, including the ability to interact with the world via physical embodiment".
How Do Humans Develop Intelligence That AI Systems Cannot Yet Replicate?
- Physical Embodiment: Humans interact with the physical world through their bodies, learning cause-and-effect relationships by manipulating objects, moving through space, and experiencing consequences firsthand. Current VLMs process images statically without this hands-on experience or ability to act on the world.
- Multi-Sensory Integration: Beyond vision, humans rely on touch, taste, smell, hearing, and proprioception (awareness of body position in space) to build a complete understanding of their environment. VLMs currently lack access to these sensory channels, limiting their ability to form comprehensive mental models of reality.
- Developmental Learning Through Physical Manipulation: Human intelligence develops gradually through childhood, with physical manipulation of objects playing a crucial role in cognitive development. Intelligence tests for children under 18 include tasks like assembling blocks and manipulating toys specifically because these physical interactions are foundational to reasoning ability.
Thompson points to a concrete example: the median human in 2024-2025 can perform everyday tasks that seem simple but require deep physical understanding. These include making a cup of coffee in an unfamiliar kitchen, assembling IKEA furniture, or navigating unexpected physical challenges. Current VLMs cannot reliably perform these tasks because they lack the embodied knowledge that humans accumulate through lived experience.
Can Vision Language Models Achieve True Intelligence Without Physical Bodies?
This question sits at the heart of current AI research debates. Some argue that embodiment is essential for genuine intelligence, while others contend that sufficiently advanced language models could achieve superhuman reasoning without ever interacting with the physical world.
Thompson's framework suggests that current VLMs and large language models (LLMs) are already approaching what he calls "proto-ASI" status, meaning they can perform at expert levels in many narrow domains. Artificial superintelligence (ASI) would match expert-level human performance across practically any field. However, current systems remain inconsistent and unreliable, producing errors and hallucinations that prevent them from being trusted as true general intelligence systems.
"The only thing they need to be ASI is more agency and automation," noted a reader in Thompson's analysis, referring to the gap between current advanced language models and true superintelligence.
Reader comment in Dr. Alan D. Thompson's AGI analysis, LifeArchitect.ai
Thompson acknowledges the counterargument by considering the example of Stephen Hawking, who made groundbreaking discoveries in theoretical physics despite severe physical limitations. However, Thompson emphasizes that Hawking had the benefit of more than two decades of full embodiment, including access to all five human senses, until ALS began to weaken his physical abilities. He also notes that all major IQ tests for children under 18 include physical object manipulation tasks, suggesting that embodied experience plays a measurable role in cognitive development.
What Role Could Robotics Play in Bridging the Gap?
Recent developments in robotics and embodied AI are beginning to address this fundamental limitation. Systems that combine vision language models with robotic bodies represent a potential bridge between pure language-based AI and truly embodied intelligence. As of July 2026, Google DeepMind has developed Gemini Robotics 2, which includes models for whole-body humanoid control with advanced multi-finger dexterity and multi-robot teamwork capabilities.
The Gemini Robotics 2 system includes three specialized models: the VLA (vision language action model) for direct control, the embodied reasoning agent for complex task planning, and an on-device version for adaptation to new robot bodies. These systems can now perform tasks requiring precise manipulation, such as tying knots and sealing ziplock bags, though success rates vary depending on the task and robot configuration.
The debate over embodiment's necessity for AGI remains unresolved, but the evidence suggests that current VLMs, despite their impressive capabilities in image analysis and reasoning, represent only one component of human-like intelligence. Until these systems gain the ability to interact with and learn from the physical world through embodied experience, they will likely remain powerful specialized tools rather than truly general intelligences capable of matching human-level understanding across all domains.