Foundation Models

Language Grounding

Language grounding is the problem of connecting words and phrases to their referents in the physical world, such as objects, spatial relations, and actions perceivable by a robot. Grounded understanding lets a system map an instruction like "put the mug left of the plate" onto specific percepts and motor commands. Approaches range from classical semantic parsing to vision-language models such as CLIP that align text and image embeddings in a shared space.

Why it matters for physical AI

Instruction-following robots only work if language reliably maps to real objects and feasible actions; grounding failures are a dominant error mode for vision-language-action models deployed in cluttered, open-world environments.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.