Google Gemini Robotics 2 brings whole-body control to robots
- Google Gemini Robotics 2 enables whole-body control and robot collaboration.
- It runs locally, adapts across robots, and includes safety controls.
Videos of humanoid robots walking, carrying objects, and completing factory or household tasks have become common online. Many of these systems, however, remain designed for defined workflows or specific types of hardware.
Google DeepMind is extending its Gemini models into this field with Gemini Robotics 2, a set of AI models for whole-body control, task planning, and coordination between different robots.
The system consists of three models with separate roles. Gemini Robotics ER 2 handles reasoning and task planning, Gemini Robotics 2 converts visual and language instructions into physical movements, and Gemini Robotics On-Device 2 runs motor-control functions locally on robotic hardware.
Gemini Robotics ER 2 can receive continuous video, audio, and text inputs before selecting the navigation, robot-control, and vision-language-action tools required for a task. Lower-level systems execute the physical movements, while ER 2 monitors progress and determines what should happen next.
Gemini controls movement and planning
Gemini Robotics 2 is Google’s latest vision-language-action model. It converts visual and language inputs into commands for humanoid robots and dual-arm systems.
The model controls movements across a humanoid robot’s legs, arms, hands, and fingers. Google said its earlier robotics models were largely limited to upper-body movements and tabletop tasks.
In one demonstration, Apptronik’s Apollo 2 humanoid was instructed to place a watering can in a green bin on the bottom shelf. The robot walked to the object, picked it up, moved to the shelves, and placed it in the requested location.
Google said the robot still needs improvements in movement speed. The demonstration combined locomotion, balance, object handling, and instruction following within a single task sequence.
Gemini Robotics ER 2 can also process continuous video and plan a subsequent action while the robot is completing its current movement. Google said the model can identify when a task begins or ends, determine whether a stage has been completed, and locate the point at which a specified event occurs.
The system also supports actions requiring more precise hand control. Gemini Robotics 2 was used to operate a five-fingered SharpaWave robotic hand with 22 degrees of freedom on the Apollo 2 platform.
The robotic hand completed tasks such as tying a knot and closing a resealable plastic bag. The model was also tested with two-fingered parallel grippers on a Franka Duo robot for packing tasks in confined spaces.
Google reported success rates ranging from 45.7% to 76.3% across three whole-body manipulation categories. The company reported rates of between 74.2% and 89.6% across three Franka Duo gripper categories, while results for multi-fingered tasks ranged from 32% to 92%.
These figures came from Google’s internal evaluations. Google said multi-fingered manipulation remains more difficult and that further work is needed to improve precision and movement speed.
Gemini Robotics ER 2 manages the system’s higher-level reasoning. It interprets instructions, examines the surrounding environment, divides a task into steps, and coordinates the movement systems used to complete them.
Google said ER 2 can manage task sequences lasting several minutes and involving hundreds of decisions. Developers can connect the model to navigation APIs, manipulator controls, vision-language-action models, and other user-defined tools.
Google demonstrated this approach using Boston Dynamics’ Spot robot. Gemini Robotics ER 2 orchestrated the robot’s navigation and manipulator APIs to retrieve objects in response to natural-language instructions.
The model can also respond when an action fails. Google said it can revise its plan, repeat a step, or select another action based on changes in the environment.
Robots divide longer workflows
The update introduces support for collaboration between different robots. Google said machines using the system can communicate, divide work, and complete workflows that one robot could not handle alone.
Google demonstrated task handoffs between Apptronik’s Apollo 2 humanoid and a Franka F3 Duo dual-arm platform. The demonstration showed how robots with different physical capabilities can contribute to separate stages of the same workflow.
The reasoning model coordinates the broader task, while separate vision-language-action models control each robot’s movements. It can also track whether one stage has been completed before another begins.
The capabilities overlap with tasks targeted by humanoid robot developers in manufacturing and logistics. Apptronik identifies material handling, line-side work, inspection, and repetitive production processes as intended applications for Apollo.
Its stated warehouse and fulfilment applications include picking, packing, and moving goods. These are proposed uses for the Apollo platform rather than confirmed commercial deployments of Gemini Robotics 2.
Apptronik also identifies hospitals and eldercare as longer-term environments for humanoid robots. Neither Google nor Apptronik has disclosed healthcare testing or deployments involving Gemini Robotics 2.
Models adapt across hardware
Gemini Robotics On-Device 2 is designed to operate directly on robotic systems. Google said the local model is intended for applications that cannot depend on continuous internet access or tolerate delays caused by remote processing.
Google said the model can be adapted to robotic platforms with different shapes, sensors, and movement capabilities. The company reported that it can adjust the system to a new dual-arm robot using several hours of data and typically fewer than 200 examples.
Differences in joints, sensors, grippers, reach, control systems, and degrees of freedom make it difficult to transfer movements directly between robotic platforms. Google’s motion-transfer approach is designed to adapt previously learned actions for machines with different mechanical configurations.
The model builds on motion-transfer methods introduced with Gemini Robotics 1.5. Google developed these methods to reduce the amount of platform-specific training required when transferring the model to another robot.
Google demonstrated the adapted model performing tasks on Dexmate, SO101, and Trossen platforms. The robots have different physical designs, sensors, and degrees of freedom.
The reported adaptation results came from Google’s internal evaluations. The company has not published the cost of collecting the required examples or integrating the model with a new commercial platform.
Safety remains a system-level issue
The release includes a robotics safety benchmark called ASIMOV-Agentic. It evaluates how an embodied reasoning model responds to unsafe instructions, uncertain situations, and tasks that cannot be completed under existing conditions.
The tests examine whether the reasoning model refuses unsafe commands, recognises when a task is not feasible, and asks a person for assistance when it cannot determine how to proceed.
In Google’s laboratory tests, ER 2 detected a person entering Apollo 2’s working area and directed the robot to place down the objects it was holding before moving into a safe pose. Google said the system resumed the task after the person left the area.
Google reported that ER 2 performed better than its previous model in internal tests involving safety constraints and human proximity. The results have not been independently reproduced.
The AI controls do not represent the complete safety system required for industrial deployment. Collaborative robots also depend on physical safeguards, emergency-stop functions, hardware controls, system integration, and workplace-specific risk assessments.
ISO 10218-1:2025 covers safety requirements for industrial robots, while ISO 10218-2:2025 addresses industrial robot applications and integration. ISO/TS 15066 provides additional guidance for collaborative applications, where safety assessments must consider the robot, its tools, the task, and the surrounding workspace.
These standards do not automatically apply to every proposed Apollo use case. Their scope excludes some service, healthcare, publicly accessible, and mobile robotic applications, which can be subject to different requirements.
An AI model’s ability to recognise a person or refuse an instruction does not establish that the complete robotic system is certified for industrial use. Deployment requirements also depend on operating speed, payload, workplace layout, hardware redundancy, and how the robot is integrated into an existing process.
Gemini Robotics ER 2 is available through the Gemini API and Google AI Studio. It is also in private preview through the Gemini Enterprise Agent Platform.
The primary vision-language-action model and the on-device version are available to selected early-access partners. Google has not disclosed production deployments, industrial operating costs, independent benchmark results, task-failure rates in uncontrolled workplaces, or measured productivity gains from multi-robot coordination.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.
Tech Wire Asia is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.