Breakthrough for Humanoid Robots: Google DeepMind Unveils Gemini Robotics 2

·
By Raisink Team

Google DeepMind researchers have unveiled a family of AI models designed to power humanoid robots. These new series, called Gemini Robotics 2, allows multiple autonomous machines to collaborate on complex tasks involving hundreds of steps.

The Gemini Robotics 2 series is built to automate chores that demand human-like dexterity and problem-solving skills. Humanoid robots usually rely on a dual-system AI architecture with two separate models working together: one crafts the high-level plan for how to perform a task, while the other translates those instructions into low-level commands for the host robot’s motors.

The new Gemini Robotics 2 series boasts an improved embodied reasoning algorithm called Gemini ER 2. This tool lets users describe tasks in natural language so humanoid robots can understand and execute complex actions even when they involve hundreds of steps or take several minutes to complete.

One key feature of ER 2 is its ability to split lengthy chores among multiple robots speeding up task completion. Users also have a tool-calling feature that gives them access to external cloud services like Google Search helping the model clarify parts it doesn’t understand when interpreting tasks.

When tasks don’t go as planned, humans and robots alike need to redo them without manual intervention. To address this challenge DeepMind’s engineers equipped ER 2 with two features: it can track task progress using footage from the host robot’s cameras and identify mistakes made along the way.

If a mistake is made ER 2 identifies the last step completed correctly and picks up where it left off allowing tasks to resume smoothly. This means robots have more time for other work and don’t get stuck on one task forever.

Gemini ER 2 also makes AI tools more efficient at determining when a task is complete enabling humanoid robots’ AI to tick off tasks faster. With this capability, robots can move on to the next step sooner and adapt in real-time situations.

Google engineers Steven Hansen and Peng Xu explained their work ‘By watching continuous video feeds robots can now track their own progress adapt if something goes wrong and know exactly when to move on to the next step.’ Their expertise has improved humanoid robot capabilities through ER 2’s integration of AI making a significant impact in this field.

After ER 2 generates a plan for task execution it sends instructions to one of two other models in the Gemini Robotics 2 series. These VLA algorithms translate action plans into low-level instructions for the host robot: the first model is called Gemini Robotics 2 and has several notable improvements over its predecessor.

Gemini Robotics 2 can control all components of a humanoid robot not just its hands allowing it to optimize the center of gravity minimizing falls in the process. This algorithm supports a broader range of robotic hands than earlier software from DeepMind giving robots more flexibility and precision when interacting with their environment.

The second VLA model is called Gemini Robotics On-Device 2 and runs directly on humanoid robots’ onboard computers. Google says this model can be adapted to new robots in just a few hours of training making the integration process smoother for developers who want to implement ER 2 in their projects.

Developers have several options when it comes to accessing ER 2: they can use Google Cloud the Gemini API or Google AI Studio. The company has also rolled out a new embodied AI safety benchmark called ASIMOV-Agentic Benchmark which evaluates human robots’ ability to avoid collisions and other risks.

This benchmark will help developers refine their models to ensure safe interactions between humans and machines. Safety is an essential consideration in robotics development, and this move by Google DeepMind sets a high standard for the industry as a whole.

Related news