Foundation Models

R3M

R3M (Reusable Representations for Robot Manipulation) is a pretrained visual encoder introduced by Nair et al. in 2022 that trains a ResNet on the Ego4D egocentric human video dataset using time-contrastive learning and video-language alignment objectives. Used as a frozen perception backbone, R3M features improved downstream imitation-learning success over training from scratch or ImageNet features across simulated and real manipulation tasks.

Why it matters for physical AI

R3M was among the first demonstrations that representations learned from human video transfer to robot control, helping establish video pre-training as a route around scarce robot data.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.