Robot Learning

Multimodal Action Distribution

A multimodal action distribution is a conditional distribution over actions with multiple distinct modes, arising in demonstration data whenever different experts, or the same expert at different times, solve identical situations differently, such as passing an obstacle on either side. Unimodal regression averages across modes and can output invalid intermediate actions; expressive policy heads based on mixture models, diffusion, flow matching, or autoregressive tokenization model the modes faithfully.

Why it matters for physical AI

Handling demonstration multimodality is a chief reason diffusion- and tokenization-based action decoders now dominate imitation learning, directly improving success on contact-rich tasks trained from diverse human data.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.