Google DeepMind launches Gemini Robotics 2 #DeepMind #GeminiRobotics2 #robotics
Google DeepMind's Gemini Robotics 2 introduces a significant advancement in humanoid robotics: real-time, generalized task execution in unstructured environments. Unlike traditional robot demos that replay memorized sequences, Gemini uses two AI models—an embodied reasoning model for high-level planning and a Vision-Language-Action (VLA) model for direct pixel-to-motor commands—to adapt to unpredictable changes, such as tying a trash bag. This system also marks the first time a single AI controls an entire humanoid robot, enabling complex full-body movements and establishing new safety benchmarks inspired by Isaac Asimov's laws.
read more
Google DeepMind has unveiled Gemini Robotics 2, a significant leap forward in humanoid robotics, moving beyond rote memorization to enable real-time, adaptable task execution in dynamic, unstructured environments. The core breakthrough lies in the robot's ability to handle highly variable tasks, exemplified by tying a trash bag, a seemingly simple action that involves continuous shape changes and unpredictable object interactions, making it impossible to pre-program. This capability is powered by a novel dual-AI model architecture.
At the higher level, an embodied reasoning model functions as the robot's 'brain.' This model observes the environment and formulates a high-level plan for a given task. It breaks down complex instructions into a sequence of actionable steps, considering the robot's physical capabilities and the current state of its surroundings. This is crucial for navigating dynamic environments and responding to unforeseen events.
These high-level steps are then passed to a Vision-Language-Action (VLA) model. This VLA model directly translates raw camera pixel data into low-level motor commands for the robot's various joints. This direct mapping eliminates the need for intermediate representations or complex inverse kinematics, allowing for more fluid and responsive movements. The VLA model is designed to handle the inherent variability of real-world objects and interactions, enabling the robot to adapt to continuous changes, such as the shifting form of a trash bag or the different sizes and textures of objects it needs to manipulate.
One of the most notable features of Gemini Robotics 2 is that it represents the first time a single AI system controls an entire humanoid robot. Previous systems often focused on controlling individual limbs or operating on a fixed tabletop. By controlling the whole body, the robot can perform complex maneuvers like crouching down to pick up objects from the floor or reaching for items on high shelves, greatly expanding its operational versatility in human environments.
Recognizing the increasing autonomy of these robots, DeepMind has also introduced a new safety benchmark named after Isaac Asimov. This benchmark is inspired by Asimov's Three Laws of Robotics, which dictate that robots must not harm humans, must obey human orders (unless conflicting with the first law), and must protect their own existence (unless conflicting with the first or second law). This initiative highlights a proactive approach to ensuring the safe and ethical deployment of advanced humanoid robots, especially as they become more integrated into daily life for tasks like folding laundry, changing light bulbs, or even powering remote outposts on Mars.