Google Introduces Gemini Robotics ER 2 with Advanced Video Understanding and Task Planning – Firstpost

Google has announced the all-new Gemini Robotics 2, a series of updates designed to give AI-powered robots greater dexterity while making them easier to control and manage. The latest release is part of Google’s broader push, announced last year, to bring its Gemini models to robotics.

According to Google, Gemini Robotics 2 enables robots to reason through every movement, tackle a broader range of tasks and continuously adapt as they work.

The update comes as the US Federal Communications Commission (FCC) moves ahead with plans to ban future sales of Chinese-made robots over security concerns. Google’s latest update aims to take robots beyond basic spatial reasoning by enabling them to think faster, make better real-time decisions and reason at the speed required in the physical world.

STORY CONTINUES BELOW THIS AD

Gemini Robotics ER 2

Google describes Gemini Robotics ER 2 as a high-level “brain” for robots. It enables robots to communicate with humans, understand the physical world and plan complex, multi-step tasks. The model can delegate motor execution to lower-level Vision-Language-Action (VLA) models while orchestrating the overall task.

ER 2 can also natively call tools such as Google Search to retrieve information or perform actions based on human prompts. According to Google, Gemini Robotics ER 2 allows robots to “think” about what comes next while simultaneously carrying out their current actions.

Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6. Robots powered by the new model can now track their own progress, adapt when something goes wrong and determine exactly when to move on to the next step. Google has also introduced multi-robot collaboration, allowing robots to work together in shared environments and complete complex tasks that would be difficult for a single robot to perform.

Advanced agentic capabilities

Gemini Robotics ER 2 stands apart because it functions as a physical AI agent. It orchestrates robotic workflows, enables self-correction and generalizes to previously unseen situations.

To build an agentic robotics system, developers can expose low-level control interfaces—such as VLA models or navigation APIs—as tools, while streaming multimodal video, audio and text directly into the model.

Because high-level reasoning in robotics depends on fast execution, Google has also integrated the Gemini Live API, featuring a bidirectional streaming endpoint optimized for latency-sensitive tasks. This enables robots to complete multi-step tasks without disruptive “stop-and-think” pauses.

STORY CONTINUES BELOW THIS AD

Robots have traditionally struggled to determine when a task has been successfully completed. ER 2 introduces a significant improvement in video understanding and progress tracking, allowing robots to verify complex tasks such as tightening a light bulb or tying a trash bag.

Improvements in spatial intelligence

Gemini Robotics ER 2 also advances Google’s core spatial reasoning capabilities, as measured across three benchmark tests.

The update improves success and failure detection by operating on raw video feeds rather than static snapshots, allowing robots to identify mid-execution failures such as spills, slips and object misalignment.

The new release also expands robots’ ability to interpret instruments beyond circular dials and sight glasses to include digital displays, linear scales, rulers and liquid thermometers. Google says the update also improves visual question answering through Gemini’s enhanced multimodal understanding.

Safety improvements

Google says Gemini Robotics ER 2 is its safest robotics model yet, delivering major improvements in Safety Instruction Following and Human Proximity benchmarks.

The model can automatically halt a humanoid robot when a person is nearby and autonomously resume work once the area is clear.

STORY CONTINUES BELOW THIS AD

Google has also introduced a new benchmark to evaluate a foundation model’s ability to safely orchestrate Vision-Language-Action (VLA) systems. The benchmark measures whether a model can enforce safety constraints, monitor its surroundings, assess the physical feasibility of actions and seek human clarification whenever necessary.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *