Foundation Models

PaLM-E

PaLM-E is an embodied multimodal language model introduced by Google in 2023 that injects continuous sensor modalities, such as images and state estimates, directly into the token stream of the PaLM language model. Its largest variant reached 562 billion parameters and handled robot task planning, visual question answering, and captioning within a single model. PaLM-E later served as one of the vision-language backbones behind the RT-2 vision-language-action model.

Why it matters for physical AI

PaLM-E was an early demonstration that web-scale vision-language pre-training transfers to embodied decision-making, shaping the now-standard recipe of building robot policies on top of pretrained multimodal backbones.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.