Vision-Language-Action Models
An overview of Vision-Language-Action (VLA) models that enable robots to understand language instructions and perform manipulation tasks.
Ecosystem Snapshot
Leading Models
Google DeepMind's most advanced vision-language-action (VLA) model, enabling whole-body control from feet to fingertips, dexterous manipulation, and multi-robot collaboration.
Google DeepMind's most capable embodied reasoning (ER) model, serving as a robot's high-level brain for environment understanding, multi-step task planning, and multi-robot collaboration, available via the Gemini API.
Google DeepMind's efficient VLA model optimized for on-device runtime, adapting to new robot embodiments with just a few hours of data.
Helix is Figure AI's proprietary VLA model for generalist humanoid control with zero-shot manipulation and multi-robot collaboration.
pi0 (pi-zero) is Physical Intelligence's generalist VLA robot foundation model for zero-shot dexterous manipulation across 8 robot types.
OpenVLA is a pioneering open-source 7B VLA model combining a pretrained VLM with action de-tokenization for zero-shot robot manipulation.
Industry Insights
This page aggregates Vision-Language-Action (VLA) models that combine internet-scale vision-language pretraining with robot control outputs. VLA models represent a paradigm shift in robotics, enabling zero-shot generalization, cross-embodiment transfer, and natural language-driven task execution.
The collection includes leading proprietary foundation models such as Gemini Robotics 2, Helix, and π0, alongside influential open-source alternatives including OpenVLA. Together these models represent the rapid evolution of Vision-Language-Action systems from research prototypes to deployable robot foundation models.