Gemini Robotics 2.0: Whole-body AI for humanoids with staged enterprise access
arstechnica.com

Gemini Robotics 2.0: Whole-body AI for humanoids with staged enterprise access

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRGoogle DeepMind launches Gemini Robotics 2.0 with whole-body AI for humanoids, improved safety, and a staged rollout with only one of three models publicly available.

Google DeepMind has launched Gemini Robotics 2.0, a three-model family that brings whole-body intelligence to humanoid robots. Only one of the three models is publicly available at launch, with the rest in private previews and early-access hardware partner programs. For AI builders and robotics teams, this release signals a shift toward physically embodied AI that can control full humanoids, coordinate multi-robot teams, and operate safely near humans.

What happened

Gemini Robotics 2.0 is built on Google's Gemini AI models and introduces a vision-language-action (VLA) approach that converts vision and language input directly into motor control. The model can control a full humanoid body from feet to fingertips, enabling dexterous tasks like manipulating objects with multi-finger hands. It also supports bi-arm robots and can coordinate multiple robots in shared spaces.

The three models are:

Availability is staged. Gemini Robotics ER 2 is available via Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The core VLA and On-Device 2 models are rolling out to early-access hardware partners.

Why AI builders should care

For teams building AI-enabled robots or automation workflows, Gemini Robotics 2.0 represents a step toward general-purpose physical AI. The whole-body intelligence means a single model can control an entire humanoid, reducing the need for task-specific controllers. The on-device AI capability promises low-latency responses, which is critical for real-time dexterity tasks.

The safety improvements in ER 2 are particularly relevant for human-robot collaboration. The model's ability to understand and stop actions when a human is nearby addresses a key barrier to deploying robots in shared workspaces. Google's release of a safety benchmark on Hugging Face also gives developers a standardized way to evaluate these capabilities.

Practical implications

If you're evaluating Gemini Robotics 2.0 for your stack, start with the ER 2 model through Google AI Studio or the Enterprise Agent Platform. Test its safety behaviors and integration with your existing hardware. The staged rollout means you'll need to work through Google's enterprise channels for the full VLA and On-Device models.

For robotics teams, the VLA approach suggests a convergence of language understanding and motor control. This could simplify the pipeline: instead of separate perception, planning, and control modules, a single model handles all three. However, the closed nature of the early-access program means you'll need to partner with Google or wait for broader availability.

Caveats

Public availability is limited. Only one of the three models is widely accessible, and the others are restricted to early-access hardware partners. Pricing, deployment options, and specific hardware compatibility details are not yet public. The staged rollout means most developers cannot yet experiment with the full capabilities.

Additionally, the claims about dexterity and task counts come from Google's own demos and announcements. Independent benchmarks are not yet available. Treat the performance numbers as preliminary until third-party evaluations emerge.

FAQs

Gemini Robotics 2.0 is a three-model family from Google DeepMind that brings whole-body intelligence to humanoid robots. Unlike the original, it controls the entire body from feet to fingertips using a vision-language-action (VLA) model, and only one of the three models is publicly available at launch.

Sources

Latest Tech News