Google is approaching the robotics race the way it once approached smartphones: build the intelligence layer, put it on as many manufacturers' machines as possible, and let the hardware ecosystem do the scaling. In his most detailed public comments on the strategy to date, Google DeepMind CEO Koray Kavukcuoglu told The Information's AI Agenda Live that the company's advantage sits in models rather than in robot hardware itself.
"We want to partner with other robotics companies, have them use this system, and in a safe and good way, inject intelligence into all robots," he said. Asked directly whether Google would build its own humanoid, Kavukcuoglu did not rule it out — but made clear the near-term focus is model capability and partnerships. The comparison to Android's early days is hard to miss: between 2007 and 2008, co-founder Sergey Brin answered the same question about phones by pointing to partnerships with T-Mobile and HTC. Google only moved into first-party hardware years later with Nexus and Pixel.
The technical centerpiece is Gemini Robotics, a version of Gemini optimized for physical control. Google's approach has been to add actions as an output modality: Gemini Robotics is a vision-language-action model that can directly drive robots, while Gemini Robotics-ER splits out the embodied-reasoning layer so that a high-level model understands space and plans actions, then calls a robot's existing control stack rather than replacing it.
That separation matters commercially. Google does not need to write low-level motor commands for every chassis ever built. With Gemini Robotics ER 2, released in July, the model adds real-time spatial reasoning, multi-step planning, video-based progress monitoring and coordination across different machines, and it is served through the Gemini API, Google AI Studio and the company's enterprise agent platform — not tied to a Google-made robot.
A Boston Dynamics Spot demonstration shows the pattern in practice. Gemini Robotics ER 2 interpreted a natural-language request, reasoned about the environment and orchestrated Spot's existing navigation and manipulator APIs to retrieve an object. Google did not have to touch Boston Dynamics' locomotion or control stack; the model sat above it and decided how existing capabilities should be combined. DeepMind is also working with Boston Dynamics on Atlas and with Apptronik on humanoid systems, with Agile Robots and more than 100 trusted testers — from enterprise automation vendors to robotics startups — in the program.
Kavukcuoglu frames the whole effort as an extension of Google's AGI work: the same multimodal architecture handling text, vision and audio is being pushed into physical control rather than having a separate model built per task. His proof point is Waymo, which he called a "very, very safe way" to demonstrate physical AI at scale. If Gemini's architecture can underpin commercial autonomous driving, the argument goes, its transferability to other physical domains becomes far more credible.
The strategy sets Google apart from Tesla and OpenAI-backed Figure, both of which are chasing a flagship humanoid. It also means Google's robotics revenue, if it comes, will look more like a platform licensing business than a hardware business. The open question is whether robot makers will accept a layer that sits above their own control stacks — and how much of the value in embodied AI ends up being captured by whoever owns the model rather than whoever owns the actuators.
Comments (0)
Log in to join the discussion
Log InNo comments yet