Foundation Models

RT-2

RT-2 is a vision-language-action model from Google DeepMind, introduced in 2023, created by co-fine-tuning pretrained vision-language models, PaLI-X and PaLM-E, on robot trajectories with actions expressed as text tokens in the model's existing vocabulary. Co-training on web data alongside robot data transferred semantic knowledge to control, yielding emergent capabilities such as reasoning about novel object categories, icons, and simple symbolic instructions, with variants up to 55 billion parameters.

Why it matters for physical AI

RT-2 defined the vision-language-action recipe now standard across the field: start from a web-pretrained multimodal backbone and treat robot actions as another language the model learns to speak.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.