Gemini Robotics ER 2: A Breakthrough in Embodied Reasoning for Robots

·

Robotics has long been hindered by the inability of machines to think and act with the same speed as humans. While spatial reasoning is essential, it’s not enough – robots need to be able to reason quickly and make decisions in real-time to effectively assist us in everyday environments. To address this challenge, Gemini Robotics has developed ER 2, a significant upgrade over its predecessor that enables robots to think fast and act even faster.

ER 2 is designed as a high-level brain for robots, allowing them to chat with humans, understand the physical world, and plan multi-step tasks. It can also hand off motor execution to lower-level vision-language-action (VLA) models and natively call tools like Google Search or any other user-defined function. This design enables the robot to think about what comes next while simultaneously performing its actions – a crucial aspect of embodied reasoning.

One of ER 2’s key features is its ability to watch continuous video feeds, track progress, adapt if something goes wrong, and know exactly when to move on to the next step. This represents a significant upgrade over ER 1.6 and enables robots to work together in shared spaces, completing complex workflows that a single robot could not do alone – a capability known as multi-robot collaboration.

ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and private preview on Gemini Enterprise Agent Platform. To help get started, examples of how to configure the model and prompt it for more useful physical AI tasks are being shared by the development team.

In robotics, high-level reasoning depends heavily on execution speed – a challenge that ER 2 addresses through its integration into the Gemini Live API. This uses a bidirectional streaming endpoint optimized for latency-sensitive tasks, resulting in fluid orchestration: ER 2 commands action models and robotics APIs to complete multi-step tasks without jarring ‘stop-and-think’ pauses.

To illustrate this capability, a demo has been built with Spot from Boston Dynamics partners – using ER 2 to orchestrate Spot APIs such as navigation and manipulator movement. The result is an interactive robot that fetches objects for you on natural language commands. This showcases the potential of ER 2 in real-world applications.

One of robotics’ hardest challenges is knowing when a task is done, which ER 2 addresses through its video understanding and progress tracking capabilities. It brings a step-change to complex tasks such as tightening light bulbs or tying trash bags – verifying that they are complete before switching to the next task. This is achieved through two foundational capabilities: progress classification and moment finding.

Progress classification refers to a robot’s ability to track progress towards task completion, quantifying it into five levels of progress (0-20%, 20-40%, 40-60%, 60-80%, 80-100%). ER 2 achieves 57.4% accuracy on this task, outperforming previous generation models and competing frontier models.

Precision moment-finding measures a model’s ability to identify the exact video frame where a critical event takes place – such as when to stop pouring coffee into a cup. ER 2 achieves significant gains in performance on this task, enabling robots to precisely switch between tasks, verify success, and suggest corrections with high accuracy.

ER 2 also enables multi-robot collaboration by allowing diverse machines to communicate via shared semantic understanding, handoff, and complete complex tasks together. This is demonstrated through the collaboration of Apptronik’s Apollo 2 and Franka F3 Duo robots using ER 2 – showcasing its potential in real-world applications.

ER 2 advances Gemini Robotics’ core spatial reasoning capability across three benchmarks: success/failure detection, general instrument reading, and enhanced spatial VQA. It consistently achieves the highest accuracy across all core capabilities, with highlights including success detection (image/video), Question Answering (ERQA), and generalized instrument reading – a significant improvement over ER 1.6.

Advancing safety for embodied intelligence is also a key aspect of ER 2, achieving significant gains on Safety Instruction Following and Human Proximity benchmarks. It successfully halts a humanoid robot when a person is nearby and autonomously resumes work once the area is clear – demonstrating its capacity to enforce safety constraints, monitor the environment, assess physical feasibility, and seek human clarification.

ER 2 outperforms ER 1.6 and other frontier models on Safety Instruction Following and Human Proximity benchmarks, marking a significant step forward in ensuring safe operation of robots in real-world environments. The development team is committed to pushing these models towards even more complex tasks to accelerate the development of helpful robots and support the robotics community.

ER 2’s capabilities have far-reaching implications for AI-generated images and their applications – particularly when it comes to using them as a tool for businesses, enabling developers to create more sophisticated physical agents that can assist humans in everyday environments. As ER 2 continues to advance, we can expect even greater potential for robots to make a meaningful impact on our lives.