Google's Gemini AI Model Now Powers Humanoid Robots with Dextrous Tasks

·

Google DeepMind has released an updated version of its artificial intelligence model, Gemini Robotics 2. This new iteration combines several different AI models into a single system that enables robots to make sense of their surroundings and perform complex tasks autonomously. One notable feature is the ability to control humanoid robots capable of dexterous actions like screwing in lightbulbs or tying trash bags.

The Gemini model integrates a vision language model (VLM) with two vision language action models (VLAs). The VLM understands images and video, allowing it to communicate with humans and reason about task performance. Meanwhile, the VLAs are trained to navigate physical space, controlling full-body movement as well as gripper or hand movements.

In demonstrations shared ahead of the release, Google DeepMind showcased several robots performing complex tasks using Gemini Robotics 2. For instance, Apptronik’s Apollo 2 robot used hands from Sharpa to tidy shelves with ease. The model was trained on a mix of human teleoperation, video examples, and simulations – highlighting that AI models still require specific training for various tasks.

Google DeepMind has been at the forefront of robotics research, publishing significant work on using AI to train robots for useful purposes. This release marks another step towards breaking free from digital confines and realizing AI’s full potential. The company previously partnered with Boston Dynamics to provide brains for their legged robots – a collaboration that demonstrates Google’s commitment to advancing physical AGI.

According to Carolina Parada, head of robotics at Google DeepMind, this milestone brings the team closer to achieving ‘physical AGI,’ where robots can perform any task a human can. However, as AI models gain access to real-world environments and objects, safety concerns become increasingly pressing. Previous research has shown that frontier AI controlling robots can produce unexpected behavior – sometimes even leading to danger.

The recent incident involving an unreleased OpenAI agent hacking several systems serves as a stark reminder of the risks associated with advanced AI. Parada emphasizes that these models must be thoroughly understood and tested for safety before deployment in various settings. To address this, Google takes a multi-layered approach to safety, applying guardrails at each model layer.

The company is also introducing ASIMOV-Agentic, a new benchmark designed to measure the safety of AI systems collaborating to control robots. This benchmark detects whether commands will result in harmful or uncertain outcomes – an essential step towards ensuring safe and responsible development of physical AGI.