Foundation Models
Gemini Robotics
Gemini Robotics is a family of vision-language-action models from Google DeepMind, introduced in 2025, that builds robot control on top of the Gemini multimodal foundation model. The initial release paired a VLA that outputs robot actions with Gemini Robotics-ER, an embodied-reasoning variant for spatial understanding, and demonstrated dexterous bimanual tasks and cross-embodiment operation; subsequent versions added on-device execution and agentic long-horizon behavior.
Why it matters for physical AI
Frontier-lab entries like Gemini Robotics test whether internet-scale multimodal pretraining transfers decisive advantages to control, and their embodied-reasoning splits illustrate one architecture for coupling semantic planning with low-level action.
Related terms
Build physical AI
Put these concepts to work on real hardware
Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.