RoboDojo Benchmark Reveals Major Gap Between Embodied AI and Human-Level Robot Performance
13 August 2026 05:50 PM
Summary: Researchers have introduced RoboDojo, a unified benchmarking platform for evaluating embodied AI and robotic manipulation across both simulated and real-world environments. Early results reveal a significant performance gap between leading AI models and human experts in complex robotic tasks.
The Multimedia Laboratory (MMLab) at the University of Hong Kong (HKU), in collaboration with researchers from nearly 20 international institutions including the University of California, Berkeley and Tsinghua University, has launched RoboDojo, a new benchmark designed to evaluate embodied AI, robotic manipulation, and robot learning systems under a unified framework.

As embodied AI becomes a key focus of next-generation robotics, the industry has faced a major challenge: the lack of standardized evaluation methods. Existing benchmarks often rely solely on simulation environments or use inconsistent hardware setups and scoring systems, making it difficult to compare results across different robotic platforms.
RoboDojo addresses this challenge by combining simulation-based testing, standardized real-robot evaluation, and policy benchmarking into a single platform. The benchmark currently includes 42 simulation tasks, 18 real-world robotic tasks, and 30 representative robot policies, measuring capabilities such as generalization, memory, precision, and long-horizon task execution.

Initial benchmark results highlight how far current robotic AI systems remain from human-level performance. The highest-performing AI model achieved success rates of only 8.8% in simulation and 12.8% in physical robot tests, while human experts reached 76.03% and 100%, respectively. The findings underscore the need for more robust and reliable AI models capable of handling complex, multi-step tasks in dynamic real-world environments.
According to the research team, RoboDojo is among the first large-scale platforms to unify simulation and standardized real-robot evaluation, providing a transparent and reproducible benchmark for the global robotics community. With strong adoption from researchers and industry developers, the platform is expected to accelerate advances in robotics benchmarking, embodied intelligence, robot foundation models, and real-world autonomous systems, helping move robotic AI from laboratory demonstrations toward practical deployment.
