Foundation Models
R3M
R3M (Reusable Representations for Robot Manipulation) is a pretrained visual encoder introduced by Nair et al. in 2022 that trains a ResNet on the Ego4D egocentric human video dataset using time-contrastive learning and video-language alignment objectives. Used as a frozen perception backbone, R3M features improved downstream imitation-learning success over training from scratch or ImageNet features across simulated and real manipulation tasks.
Why it matters for physical AI
R3M was among the first demonstrations that representations learned from human video transfer to robot control, helping establish video pre-training as a route around scarce robot data.
Build physical AI
Put these concepts to work on real hardware
Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.