Back to Search
G

Gemini Robotics 2

ModelActive

Gemini Robotics 2 is Google DeepMind's most advanced vision-language-action (VLA) model, introduced on July 30, 2026 as the action-execution core of the Gemini Robotics 2 generation. It converts vision and language input directly into motor control, enabling robots to take action. It is the successor of Gemini Robotics 1.5 and is capable of controlling full humanoids, from feet to fingertips, as well as bi-arm robots. Core capability upgrade: whole-body control. While previous models focused on upper-body, table-top tasks, Gemini Robotics 2 expands physical AI into whole-body motions — walking, crouching, stretching, balancing, and manipulating objects at the same time, e.g. cleaning up a cluttered room. In the official demo it drives Apptronik's Apollo 2 humanoid to execute the instruction "put the watering can into the green bin in the bottom shelf", combining walking, reaching, and precise placement. Dexterity: the model controls five-fingered 22-DoF hands (e.g. the SharpaWave hand on Apollo 2) for delicate actions like tying knots, sealing ziplock bags and tightening light bulbs, and also operates standard two-fingered parallel grippers (e.g. tight packing on the Franka Duo platform). Ecosystem: Gemini Robotics 2 works with Gemini Robotics ER 2 as the high-level "brain" (task planning, environment understanding) while the VLA handles motor execution, and with Gemini Robotics On-Device 2 as the local, low-latency variant. The release is accompanied by the ASIMOV-Agentic safety benchmark and the Gemini Robotics 2 Safety Technical Report. Google DeepMind describes the release as a milestone toward solving AGI in the physical world.

Details

Updated:7/31/2026
open sourcefalse
release date2026-07-30
paper urlhttps://storage.googleapis.com/deepmind-media/gemini-robotics/Gemini-Robotics-2-Safety.pdf
model familyGemini Robotics

Tags

vlavision-language-actionwhole-body-controldexterous-manipulationhumanoidgoogle-deepmind