Google DeepMind Brings Gemini Robotics 2 to Humanoids, Enhancing Automation Capabilities

·
By Raisink Team

A significant development in the field of robotics has been announced by Google DeepMind. The company has unveiled Gemini Robotics 2, a physical AI model that brings full-body autonomous functionality, multi-step execution, real language communication, and added safety features to humanoids and other bi-manual embodiments.

The new model is available in three flavors: the standard Gemini Robotics 2 VLA (vision language action mod), Gemini Robotics ER 2, a VLM (vision language model) for embodied reasoning/human to robot communication, and Gemini Robotics On-Device 2, an edge-based model. All three are currently available in early-access form.

According to DeepMind’s VP and Head of Robotics, Carolina Parada, the models have been tested at Robot Park, a massive data collection facility in Austin, TX. The company has also showcased their capabilities on Apptronik’s Apollo 2 humanoid robot, which is capable of full-body autonomous functions and advanced reasoning in real time.

The software allows Apptronik’s Apollo 2 to perform tasks such as picking up objects and placing them in specific locations. For example, when controlling the robot, users can ask it to ‘put the watering can into the green bin in the bottom shelf.’ The robot then processes the instruction and carries out the task with precision.

Parada notes that full-body movement still has a way to go, particularly when it comes to speed. However, the new model marks a significant improvement over its predecessor, enabling robots to execute complex tasks more efficiently.

The ER 2 variant also allows systems to execute complex, multi-step tasks and correct themselves if they get stuck somewhere in the middle. This feature is crucial for automating industrial workflows, where tasks often involve multiple steps and require precise execution.

In addition to its capabilities on humanoids, Gemini Robotics 2 has been used with other robots, including those equipped with Wave hands from Sharpa for highly dexterous tasks such as knot tying and bag sealing. The system also operates with the two-fingered Franka Duo system.

Another significant feature of Gemini Robotics 2 is its ability to facilitate multi-robot communication. This allows different embodiments to communicate and coordinate in uniform workflows, which can be a game-changer for industrial settings where multiple robots are involved.

The safety features of Gemini Robotics 2 have also been enhanced with the introduction of ASIMOV-Agentic, a benchmark for agentic safety orchestration and uncertainty resolution. This feature enables robots to reject commands involving potentially unsafe tool usage and call for human assistance if they’re unsure about completing a task.

Overall, Google DeepMind’s Gemini Robotics 2 represents a significant step forward in the development of humanoid robots and their applications in automation. As companies continue to explore ways to automate industrial workflows, this new model is likely to play an important role.

Related news