Back to Search
R

Rho

ModelActive

Rho is a family of open-weights vision-language-action models for bimanual manipulation introduced in "Rho: A Foundation for Efficiently Adaptable VLA Models" (arXiv 2609.38164, submitted 2026-09-29). The backbone is distilled from Microsoft's Phi-family VLM, coupled to a flow-matching action expert generating continuous action chunks, and pretrained in end-effector-pose action space on a multi-task, multi-embodiment data mixture dominated by dual-arm robot demonstrations plus robotics-relevant web VQA data. The training recipe combines physical grounding, horizontal embodiment midtraining, and vertical task adaptation via regularized progressive adaptation, yielding embodiment-specific variants Rho-YAM-Box (I 2 RT YAM Box), Rho-UR-AI-Trainer (Universal Robots AI Trainer), and Rho-FR3-Duo (Franka FR3 Duo). Each variant's flow module includes an optional internal latent policy that learns from corrective feedback to select observation-conditioned noise inputs for the frozen action expert, enabling online adaptation with as few as 15 corrected episodes. The paper's studies were performed in the RoboTwin, LIBERO, and RoboEval simulation environments, and the authors state they release the base Rho model and the embodiment-specific checkpoints via the microsoft/rho HuggingFace collection ("a family of physical AI models for robotic manipulation"), positioning Rho as both a general-purpose manipulation model and a practical foundation for adaptation.

Details

Updated:10/1/2026
open sourcetrue
release date2026-09-29
paper urlhttps://arxiv.org/abs/2609.38164
huggingface urlhttps://huggingface.co/collections/microsoft/rho

Tags

vlabimanual-manipulationflow-matchingmicrosoftphiopen-weightsembodied-ai

Relationships

Sources

Rho: A Foundation for Efficiently Adaptable VLA Models
paper
Visit
microsoft/rho HuggingFace collection
huggingface
Visit

Appears In

No related landscapes.