ASIMOV-Agentic
BenchmarkActiveASIMOV-Agentic is a safety benchmark introduced by Google DeepMind on July 30, 2026 alongside Gemini Robotics 2, designed to evaluate agentic safety orchestration and uncertainty resolution of embodied reasoning agents. It measures, for example, the embodied reasoning agent's ability to refuse unsafe tool calls from a VLA, to predict whether a task is possible, and to proactively request human intervention when uncertain. Two layers, one release: the benchmark is built on top of a multimodal dataset (the official google/asimov_agentic dataset on Hugging Face, CC-BY-4.0, containing image, text and video material). When referring to its evaluation protocol, scoring, or leaderboard, ASIMOV-Agentic denotes the benchmark; when referring to its test items or fine-tuning corpus, it denotes the underlying multimodal dataset. Context: the benchmark targets foundation models acting as safe VLA orchestrators — enforcing safety constraints, monitoring the environment, assessing physical feasibility, and seeking human clarification. Details are documented in the Gemini Robotics 2 Safety Technical Report. It is a companion artifact to the Gemini Robotics ER 2 model, which Google DeepMind reports as its safest robotics model to date.
Details
No structured details available.