Data & Benchmarks

CALVIN

CALVIN (Composing Actions from Language and Vision) is a simulated benchmark for long-horizon, language-conditioned manipulation, introduced by Mees et al. in 2022. A Franka arm in four tabletop environments must complete chains of up to five sequential instructions such as opening drawers, pushing blocks, and toggling lights, using onboard RGB observations. Its evaluation protocol measures how many consecutive subtasks a policy completes, testing both instruction grounding and long-horizon consistency.

Why it matters for physical AI

CALVIN remains a standard yardstick for language-conditioned policies and hierarchical vision-language-action systems, exposing failure modes in instruction chaining that single-task benchmarks miss.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.