Robot Learning

Learning from Observation

Learning from observation is imitation learning from state or video sequences that lack action labels, requiring the agent to infer how to reproduce observed behavior without knowing the exact commands the demonstrator issued. Techniques include learning inverse dynamics models to recover pseudo-actions, matching state-occupancy distributions as in GAIfO, and extracting latent actions from raw video. Human videos are the most important application domain.

Why it matters for physical AI

Internet-scale human video contains far more manipulation experience than all robot datasets combined; action-free imitation is the main route to tapping it for pretraining generalist robot policies.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.