Gemini Robotics 2: Google DeepMind’s Humanoid Robot AI Explained

Google DeepMind released Gemini Robotics 2 on 30 July 2026, a vision-language-action model that controls an entire humanoid robot: walking, balancing, crouching, and handling objects with five-fingered hands. Apptronik’s Apollo 2 is the demonstration platform.

The previous generation drove a robot’s upper body. This one runs legs, torso, arms and fingers under a single learned policy, which is what DeepMind means by whole-body control.

The strategic story is bigger than the demo reel. Google is not trying to win the humanoid hardware race. It is positioning Gemini as the intelligence layer inside robots that other companies build.

What Gemini Robotics 2 actually is

DeepMind shipped three models, and they do different jobs.

Gemini Robotics 2 is the vision-language-action model. It turns camera input and spoken instructions into motor commands, and it can drive full humanoids as well as two-armed robots using either multi-finger hands or conventional parallel grippers.

Gemini Robotics ER 2 is the embodied reasoning model, built on Gemini 3.5 Flash. DeepMind describes it as the robot’s high-level brain. It holds a conversation, plans multi-step tasks, tracks its own progress, calls tools, and coordinates more than one robot on a shared job. It accepts interleaved text, images, video and audio with a 128k context window.

Gemini Robotics On-Device 2 runs locally, without a network connection. It is built on Gemini Robotics 1.5 and Google’s Gemma models, and it adapts to an unfamiliar robot body in a few hours using fewer than 200 example demonstrations.

What the demonstrations show

The clearest example is mundane on purpose. Told to put the watering can into the green bin on the bottom shelf, Apollo 2 walks across the room, finds the can, picks it up, crouches, and places it where it was asked. Nobody scripted the sequence of movements.

Other demonstrations cover unscrewing a light bulb, tying a knot in a rubbish bag, sealing a ziplock bag, sweeping debris into a dustpan, and packing items tightly using grippers. Some run for several minutes. DeepMind also shows two robots working the same task, coordinated by the ER model.

The hardware on show includes Apollo 2 fitted with SharpaWave and Inspire hands, a Franka Duo with a Robotiq gripper, and smaller research platforms from Dexmate and Trossen plus the open-source SO101 arm.

The success rates Google published

This is the part worth reading twice. DeepMind published task-level numbers, and they are candid about how uneven performance still is.

Task Robot and hand Success rate
Unscrew a light bulb Apollo 2, SharpaWave 92%
Tie a rubbish bag Apollo 2, SharpaWave 44%
Seal a ziplock bag Apollo 2, SharpaWave 40%
Screw a light bulb in Apollo 2, SharpaWave 36%
Sweep into a dustpan Apollo 2, SharpaWave 32%
Pick up from a shelf Apollo 2, Inspire 76.3%
Pick up from a table Apollo 2, Inspire 68.4%
Pick up from the floor Apollo 2, Inspire 45.7%
Precise insertion Franka Duo, gripper 89.6%
Diverse tool kitting Franka Duo, gripper 78.9%
General pick and place Franka Duo, gripper 74.2%

The light bulb pair is the tell. Taking a bulb out succeeds 92% of the time. Putting one back in succeeds 36% of the time. Removal tolerates rough alignment, insertion does not. That gap is the current state of robot dexterity in a single line.

Why this matters more than the demo

Humanoid hardware is fragmenting. Apptronik, Figure, Tesla, Unitree, Boston Dynamics, AgiBot and a dozen others are all building bodies. Foundation models are consolidating instead, because training them takes data and compute that most robot makers cannot fund alone.

That sets up an Android-shaped possibility: many robot bodies, a small number of intelligence suppliers underneath them. Google already demonstrates Gemini across humanoids, two-armed systems, multi-finger hands and grippers, and it signed an AI partnership with Boston Dynamics in January 2026 covering the next-generation Atlas.

That reading is an interpretation of Google’s multi-platform strategy, not a proven commercial outcome. No licensing terms have been published and the whole-body model is not on sale.

The reality check

A demonstration is not a deployment. Several things are worth holding onto before treating this as solved.

  • You cannot buy it. ER 2 is available in Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform. The whole-body model and the on-device model go to early-access partners through a trusted tester programme.
  • Dexterity is inconsistent. Three of the five hand tasks DeepMind published complete less than half the time.
  • Robots are still slow. DeepMind says movement speed is an area where there is more to advance.
  • The on-device model has stated limits. DeepMind says it struggles to generalise to out-of-distribution tasks and to control robots with a high number of degrees of freedom.
  • None of these figures describe a real building. They say nothing about uptime, intervention rate, safety incidents or cost per completed task.

On safety, DeepMind introduced ASIMOV-Agentic, a benchmark for agentic safety orchestration and uncertainty resolution, and showed the model detecting a nearby human and bringing the robot to a safe stop. That is real engineering. It is also a benchmark, not a certification or an insurer’s sign-off.

What it means for humanoid manufacturers

Apptronik gains the most immediately. Apollo is the platform every clip is filmed on, and Apptronik runs a 90,000 square foot facility called Robot Park in Austin, opened in July 2026, that collects training data feeding Gemini Robotics.

Boston Dynamics is the other named beneficiary. Its January 2026 partnership puts Gemini Robotics models into the new Atlas perception and reasoning stack, and DeepMind is one of only two customers taking 2026 Atlas units.

Figure and Tesla sit on the opposite side of the bet. Both build their own hardware and their own models. If Google’s layer wins, vertical integration starts to look expensive. If it does not, they keep the whole margin. Our Figure 03 versus Apollo 2 comparison covers how differently those two software stacks are being built.

Smaller manufacturers get the most interesting option of all: license intelligence rather than fund a frontier model. That only becomes real when Google publishes commercial terms.

What buyers should take from this

Gemini Robotics 2 is good evidence that general robot intelligence is improving faster than most operations teams expect. It is not evidence that you can buy a humanoid, install Gemini, and automate arbitrary work.

If you are evaluating a humanoid this year, the questions have not changed. Ask for task-level completion rates on your task, not on a benchmark. Ask for intervention rate per shift. Ask for safety documentation and the deployment support model. Ask for total installed cost rather than unit price. A 92% score on a light bulb tells you nothing about your production line.

For where the current machines actually stand, see our ranking of the best humanoid robots, and follow the humanoid robot news tracker for what lands next.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *