Robot Learning

Reward Model

A reward model is a learned function that predicts reward or task success from observations, replacing or augmenting hand-specified reward signals. Reward models are trained from human preference comparisons, as in RLHF, from success-labeled examples, or increasingly by prompting vision-language models to judge task completion from images. On robots they also serve as autonomous success detectors that let systems label their own experience.

Why it matters for physical AI

Scalable robot learning needs supervision that does not require a human watching every trial, and learned reward and success models are the main candidates for closing that autonomous feedback loop.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.