Robot Learning

Visual Imitation

Visual imitation is imitation learning in which the policy consumes raw camera images rather than privileged state, and, in a broader sense, learning behaviors from visual observation of another agent, including humans, without access to expert actions. The first form dominates modern manipulation through methods like ACT and Diffusion Policy; the second, learning from human video, must additionally bridge the embodiment gap between demonstrator and robot.

Why it matters for physical AI

Cameras are the most scalable sensor for collecting demonstrations, and progress in visual imitation determines how directly the vast supply of human manipulation video can be converted into robot skills.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.