Perception
Monocular Depth Estimation
Monocular depth estimation is the prediction of per-pixel scene depth from a single RGB image, an ill-posed problem solved by learning statistical regularities of scene geometry from data. Foundation models such as MiDaS and Depth Anything, trained on large mixed datasets, produce robust relative depth zero-shot, while metric variants recover absolute scale. Self-supervised training from video with photometric consistency avoids the need for ground-truth depth.
Why it matters for physical AI
Extracting geometry from cheap RGB cameras removes the need for depth sensors in many manipulation and navigation stacks, and dense depth priors strengthen grasping, obstacle avoidance, and 3D scene representations.
Build physical AI
Put these concepts to work on real hardware
Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.