Back to Search
H

Hy-Embodied-VLM-1.0

ModelActive

Hy-Embodied-VLM-1.0 is an embodied foundation model developed by Tencent Hunyuan, built on the Hy3-A3B language backbone and Hy-ViT2 vision encoder with Mixture-of-Experts architecture. It defines an action-centric capability taxonomy across three dimensions: Action-Relevant State Understanding, Action-Transition Reasoning, and Sequential and Adaptive Reasoning. Evaluated on 38 benchmarks covering embodied perception, physical-world understanding, and embodied reasoning, it achieves best performance on 19 benchmarks and substantially outperforms similar-sized competitors including Qwen3.6-A3B and Cosmos 3. Compared with Hy-Embodied-0.5 MoT-2B, it improves average performance by 8.4% while activating only 3B parameters. Open-sourced under Apache-2.0 license.

Details

Updated:7/30/2026
open sourcetrue
release date2026-07-14
github urlhttps://github.com/Tencent-Hunyuan/HY-Embodied
paper urlhttps://arxiv.org/abs/2607.12894
model familyHY-Embodied

Tags

VLAEmbodied Foundation ModelMoETencentPhysical-World AgentsOpen-SourceApache-2.0

Relationships

Sources

HY-Embodied GitHub
github
Visit
HY-Embodied Tech Report
paper
Visit

Appears In

No related landscapes.