What it is
Gemini Robotics is a vision-language-action model built on Gemini 2.0 that directly outputs robot control, performing smooth, reactive manipulation while staying robust to changes in object type and position and following open-vocabulary instructions. With fine-tuning it takes on long-horizon dexterous tasks, learns new short-horizon skills from as few as 100 demonstrations, and adapts to entirely new robot bodies. It rests on a second model, Gemini Robotics-ER, which adds embodied spatial and temporal reasoning (object detection, pointing, grasp and trajectory prediction, 3D bounding boxes).
Why it matters
It is a leading demonstration that a general multimodal model can carry its world knowledge into physical control, roughly doubling generalization benchmarks and helping define the vision-language-action paradigm for robotics. The two-model split (a reasoning backbone plus an action model) is a template others are now following.
Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.
Filed underrobotics, vision-language-action, embodied AI, manipulation, Gemini