Google DeepMind releases Gemini Robotics 2, a robotics system that combines multiple AI models to let humanoid and other robots perceive, understand language, and act in physical environments. According to reporting, the system integrates a vision-language model that interprets images and video and can interact with people, along with vision-language-action models trained to manage movement and control both full-body motion and end-effector actions such as grippers or hands.

Demonstrations shared before the release show robots completing a range of practical tasks using the unified model, including activities like tidying shelves, tying bags, and replacing lightbulbs. In one example, Apptronik’s Apollo 2 robot uses hands from Sharpa to organize shelves. Google says the model is trained using a combination of human teleoperation, video examples, and simulations, and that broad, general performance across complex tasks still requires task-specific training.

Google also describes safety work as a multi-layered approach, applying “guardrails” across different model components. It is also introducing the ASIMOV-Agentic benchmark to evaluate safety in scenarios involving multiple AI systems collaborating to control a robot.