Google brings Gemini Robotics ER 2 to developers through Gemini API

1 week ago 7



Google has launched Gemini Robotics ER 2, an embodied reasoning model that helps robots understand their surroundings, communicate with people, and complete complex physical tasks in real time. The model acts as a high level brain, planning tasks before passing physical execution to a vision language action model or another robot control system. Gemini Robotics ER 2 can process continuous video, audio, and text while using tools such as Google Search, navigation systems, and developer defined functions. It can plan its next step while a robot is still acting, reducing pauses between decisions. The model is available through the Gemini API and Google AI Studio, with a private preview offered through the Gemini Enterprise Agent Platform. Gemini Robotics ER 2 can track task progress, identify mistakes, and adjust actions without restarting an entire workflow. Google said the model achieved 57.4% accuracy in progress classification tests, which measure how well it estimates how much of a task has been completed. It also reached 91.3% accuracy in moment finding, which tests whether a robot can identify the exact point when an event occurs, such as when to stop pouring liquid. Google rep...

Read Entire Article