Hy-Embodied-VLM-1.0
ModelActiveHy-Embodied-VLM-1.0 is an embodied foundation model developed by Tencent Hunyuan, built on the Hy3-A3B language backbone and Hy-ViT2 vision encoder with Mixture-of-Experts architecture. It defines an action-centric capability taxonomy across three dimensions: Action-Relevant State Understanding, Action-Transition Reasoning, and Sequential and Adaptive Reasoning. Evaluated on 38 benchmarks covering embodied perception, physical-world understanding, and embodied reasoning, it achieves best performance on 19 benchmarks and substantially outperforms similar-sized competitors including Qwen3.6-A3B and Cosmos 3. Compared with Hy-Embodied-0.5 MoT-2B, it improves average performance by 8.4% while activating only 3B parameters. Open-sourced under Apache-2.0 license.