Foundation Models

SmolVLA

SmolVLA is a compact open-source vision-language-action model released by Hugging Face in 2025, with roughly 450 million parameters, designed to run on consumer hardware including single GPUs and CPUs. It pairs a pruned SmolVLM backbone with a flow-matching action expert, was trained on community-contributed LeRobot datasets rather than proprietary fleet data, and supports asynchronous inference that decouples action prediction from execution. Despite its size, it reported performance competitive with much larger VLAs on benchmark manipulation tasks.

Why it matters for physical AI

Small, openly trained VLAs democratize robot foundation models, proving that community data and modest compute can yield deployable policies outside industrial labs.

Build physical AI

Put these concepts to work on real hardware

Axol is a dual-arm robot built for physical AI — teleoperate it, collect demonstrations, and deploy learned policies out of the box.