Foundation Models
Video Prediction Model
A video prediction model is a generative model that forecasts future video frames conditioned on past frames and, optionally, on actions or language. Action-conditioned video prediction enabled early visual model-predictive control in the Visual Foresight line of work, and modern large-scale video generators are increasingly treated as learned world simulators from which robot policies can plan or be distilled.
Why it matters for physical AI
A model that predicts how scenes evolve under actions is effectively a physics simulator learned from data, offering a path to planning and policy training that scales with web video rather than hand-built simulation.
Build physical AI
Put these concepts to work on real hardware
Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.