The number that matters is 22: the degrees of freedom in the five-fingered SharpaWave hand that Google's new robotics model can now drive on a full humanoid body. In two blog posts published July 30, Google DeepMind introduced Gemini Robotics 2, its first vision-language-action model to control an entire humanoid, in its words 'from feet to fingertips,' rather than the tabletop upper-body tasks of previous versions. Demonstrated on Apptronik's Apollo 2, the robot walks to a table, grasps a watering can, carries it across the room and shelves it, then ties knots and seals a ziplock bag with the articulated hand; a Franka Duo platform does tight packing with two-fingered grippers.
The second model is the one builders can actually touch. Gemini Robotics ER 2 is the planning brain: it sequences multi-step tasks lasting several minutes and involving what Google calls hundreds of decisions, self-corrects failed steps, calls tools like Google Search natively, and plugs into the Gemini Live API for bidirectional streaming. It is publicly available starting today through the Gemini API and Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform. Google's own evals claim 57.4% accuracy on progress classification, 91.3% on moment-finding with 0.96 seconds of mean absolute distance, at four times the execution speed of much larger models, and it reads instruments across 10 types, from digital displays to liquid thermometers.
A third model, Gemini Robotics On-Device 2, runs locally and adapts to what Google calls completely new robot bodies in a few hours, typically with less than 200 examples, demoed on Dexmate, SO101 and Trossen hardware. The multi-robot demos pair Apollo 2 with a Franka F3 Duo, and a separate one has ER 2 orchestrating Boston Dynamics Spot's APIs to fetch a snack on a natural-language command, with the code on GitHub. Named partners are Apptronik, Boston Dynamics and Agile Robots. On safety, Google published a new ASIMOV-Agentic benchmark for unsafe tool calls and human intervention, and says ER 2 halts a humanoid when a person approaches and resumes when the path clears.
The caveats are structural. Every benchmark cited above is a Google-run evaluation, and the demonstrations shown to journalists, WIRED included, were videos shared ahead of the release rather than live runs. The availability split also cuts the announcement down to size: of the three models, only ER 2 is generally reachable, while Robotics 2 and On-Device 2 remain limited to early-access partners. Carolina Parada, DeepMind's head of robotics, frames the release as a milestone toward 'physical AGI.' The quieter reading is more useful for builders: robot control is becoming a feature of the same API bill as text generation, and the open question is shifting from whether the model can move the arm to who is allowed to run it, and under what safety case.
