Perception
Depth Anything
Depth Anything is a family of monocular depth estimation foundation models, first released in 2024, trained on a combination of labeled datasets and tens of millions of unlabeled images via large-scale pseudo-labeling. It produces robust relative depth across diverse scenes, with fine-tuned variants providing metric depth. Depth Anything V2 improved detail and robustness using synthetic labels and larger teacher models.
Why it matters for physical AI
Reliable depth from a single RGB camera reduces sensor cost and calibration burden on robot platforms. Foundation-model depth priors increasingly feed grasping, navigation, and 3D scene reconstruction pipelines without per-robot training.
Build physical AI
Put these concepts to work on real hardware
Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.