Google's New Robots Can Now Walk, Grab, and Collaborate: Here's What Changes
Google DeepMind has unveiled Gemini Robotics 2, a major leap forward in robotic intelligence that enables robots to control their entire bodies, perform delicate manipulation tasks, and collaborate with other robots to complete complex real-world jobs. The new system moves beyond pre-programmed, single-task machines to create adaptable robots that can reason through their movements, learn from experience, and transfer skills across different robot bodies in just hours.
What Makes Gemini Robotics 2 Different From Earlier Robot AI?
For decades, most robots have been stuck doing the same narrow, repetitive tasks over and over. They follow pre-written instructions or require a human operator to control them remotely. Gemini Robotics 2 changes this fundamental limitation by giving robots the ability to think, act, and adapt intelligently to unpredictable environments. The system combines three specialized models working together to create what Google calls "whole-body intelligence".
The breakthrough centers on a vision-language-action model, or VLA, which converts what a robot sees and hears into physical movements. Think of it as teaching a robot to understand instructions like "put the watering can into the green bin on the bottom shelf" and then actually execute that multi-step task by walking to the table, picking up the object, navigating to the shelves, and placing it precisely in the right location.
How Do These Three Robot Models Work Together?
- Gemini Robotics 2 (VLA): The primary action model that controls full humanoid robots from feet to fingertips, enabling whole-body movement and dexterous manipulation with hands or grippers across different robot platforms.
- Gemini Robotics ER 2 (Embodied Reasoning): Acts as the robot's high-level brain, processing human instructions, understanding the physical world, planning multi-step tasks lasting several minutes, and enabling robots to communicate with humans and work together as a team.
- Gemini Robotics On-Device 2 (Efficient VLA): A lightweight version optimized to run directly on robot hardware without internet connectivity, capable of adapting to completely new robot bodies with just a few hours of training data and fewer than 200 examples.
This three-layer approach solves a critical problem in robotics: transferring learned skills from one robot body to another has historically been incredibly difficult. Gemini Robotics 2 can now adapt to new bi-arm robot embodiments with drastically different shapes, sensors, and degrees of freedom, demonstrating this flexibility across platforms like the Dexmate, SO101, and Trossen robots.
What Physical Tasks Can These Robots Actually Perform?
The new system unlocks capabilities that were previously out of reach for AI-controlled robots. Humanoids can now walk, crouch, stretch, and manipulate objects to clean up cluttered rooms. The five-fingered, 22 degree-of-freedom SharpaWave hand on the Apptronik Apollo 2 robot can perform delicate actions like tying knots or sealing ziplock bags. Standard two-fingered parallel grippers can handle complex dexterous tasks such as tight packing.
The embodied reasoning model enables robots to execute longer task sequences lasting several minutes and involving hundreds of individual decisions. Crucially, robots can now self-correct if a step fails and generalize to novel situations and goals they haven't encountered before. Google also introduced multi-robot collaboration, allowing different types of robots to communicate and work together to solve complex workflows that a single robot could not complete alone.
How Can Developers and Organizations Access These Models?
- Gemini Robotics ER 2 Availability: The embodied reasoning model is now available on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform for organizations building custom robotic solutions.
- VLA and On-Device Models: The vision-language-action models are available to early-access partners, with Google providing developer documentation on how to integrate these models with specific robot hardware platforms.
- Developer Support: Google has published detailed guidance on its Developer blog explaining how to bring these models to custom robotic hardware and adapt them to new embodiments.
The on-device model is particularly significant for real-world deployment because many robotic applications need to operate without network latency or internet connectivity. By running locally on robot hardware, the system can respond instantly to environmental changes and operate in locations where cloud connectivity is unreliable or unavailable.
What Safety Measures Are Built Into These Robots?
As robots gain more physical capabilities, Google has prioritized safety at every level. The company introduced ASIMOV-Agentic, a new benchmark specifically designed to measure agentic safety orchestration and uncertainty resolution. This benchmark evaluates whether the embodied reasoning agent can refuse unsafe tool calls from the vision-language-action model and handle real-world uncertainty appropriately.
Google's approach combines traditional physical safety measures with robust AI safety frameworks. The system is designed to navigate the uncertainty of the real world while safely collaborating alongside humans. This multi-layered safety strategy reflects the company's commitment to ensuring that as robots become more capable, they remain trustworthy and aligned with human values.
Why Does This Matter for the Future of Robotics?
The world is built for human movements. Humans reach, bend, and balance in tight, cluttered spaces that most robots have struggled to navigate. Gemini Robotics 2 expands physical AI into whole-body motions for the first time, enabling robots to operate in real homes and workplaces rather than just controlled laboratory environments. While the robots still need to advance in movement speed and precision to match human-level dexterity, this represents a fundamental shift toward machines that can truly adapt to unpredictable, real-world conditions.
The ability to transfer skills across different robot bodies in just hours, rather than weeks or months, dramatically accelerates the pace at which robotic capabilities can be deployed at scale. This means that breakthroughs in one robot platform can quickly benefit many others, creating a compounding effect in robotic progress across the industry.