Illustration for: Full-Body Robot Control Is the New AI Battleground

Full-Body Robot Control Is the New AI Battleground

Google DeepMind's Gemini Robotics 2, which extends full-body control to an entire humanoid robot rather than just its upper body, marks a shift in AI-for-robotics competition from narrow manipulation tasks to whole-body, video-native control.

By the Numbers

Gemini Robotics 2
Model
Full-body control
New capability
Video-native
Input type
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Gemini Robotics 2's leap from upper-body-only to full-body control -- walking, crouching, and manipulating objects from the same model -- collapses what used to require separate locomotion and manipulation systems into one video-native model

2

The move puts Google DeepMind in more direct competition with humanoid-robot-focused AI efforts at Tesla (Optimus), Figure, and China's Unitree, all racing to prove a single foundation model can generalize across a robot's entire body rather than needing task-specific tuning

3

Video-native input -- learning directly from video rather than requiring specialized robot-specific sensor data -- is the more consequential technical bet, since it means the same architecture could plausibly scale using the kind of data that trains ordinary vision-language models

4

For robotics investors, whichever lab's foundation model becomes the default full-body control layer captures outsized value across every humanoid-robot hardware company built on top of it, mirroring how a handful of LLMs now power most AI applications

TC

The VC Read · Trace's Take

Trace Cohen

Collapsing locomotion and manipulation into one video-native model is the actual news here, not the demo reel of a robot crouching. Whoever's foundation model becomes the default "brain" for humanoid robots captures value across every hardware company that licenses it -- which means the real investment question for robotics right now isn't which humanoid hardware company to back, it's which model layer they're all going to end up standardized on.

Analysis

Google DeepMind's Gemini Robotics 2 extends the company's robotics AI model to full-body control of a humanoid robot -- walking, crouching, and manipulating objects -- where its predecessor handled only upper-body manipulation tasks. The model is video-native, meaning it learns to control a robot's movements from video input rather than requiring specialized robot-specific sensor data.

That's a meaningful technical jump. Most humanoid-robot AI to date has stitched together separate systems for locomotion and manipulation; a single model handling both end-to-end is closer to how the large language model wave collapsed dozens of narrow NLP tasks into one general-purpose architecture. It puts Google DeepMind in more direct competition with Tesla's Optimus program, Figure, and China's Unitree, all racing to prove a single foundation model can generalize across an entire robot's body.

For robotics investors, the stakes are about which layer captures value: if a small number of foundation models end up powering most humanoid robots the way a small number of LLMs now power most AI applications, the hardware companies building on top of those models may end up with less differentiated economics than the model layer itself. What to watch: whether Gemini Robotics 2 gets licensed to third-party humanoid-robot hardware makers, and how it compares head-to-head against Tesla's and Figure's own foundation models on real-world tasks.

ShareXLinkedInEmail

More on

Google

Key Sources

2 sources

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.