|
Work item:
|
F.BA-EAI
|
|
Subject/title:
|
Framework for benchmarking and assessment of embodied artificial intelligence systems
|
|
Status:
|
Under study
|
|
Approval process:
|
AAP
|
|
Type of work item:
|
Recommendation
|
|
Version:
|
New
|
|
Equivalent number:
|
-
|
|
Timing:
|
2027-12 (Medium priority)
|
|
Liaison:
|
-
|
|
Supporting members:
|
HUAWEI, Peng Cheng National Laboratory , Ant Group Co., Ltd,CAICT
|
|
Summary:
|
Embodied Artificial Intelligence (EAI) integrates intelligence with physical entities capable of perceiving, reasoning, learning, interacting, and acting within real-world environments. This Recommendation establishes a unified framework for the benchmarking and assessment of EAI systems, specifying the general benchmarking framework, the benchmarking methods, and the associated evaluation metrics.
The benchmarking framework is organized as a workflow of five stages: determination of benchmark dimensions, test environment setup, test case configuration, test execution, and metric calculation. The benchmark dimensions comprise basic capability evaluation, generalization evaluation, robustness evaluation, and advanced capability evaluation. Test environments includes simulation, real-world, or hybrid. Benchmark tasks are organized as atomic actions, short-horizon tasks, and long-horizon tasks. The evaluation metrics include task success rate, task completion rate, task execution efficiency, human intervention, task success rate decay, and average task energy consumption. It supports objective capability measurement, technology verification, and performance comparison, while promoting transparency, interoperability, and sustainable development of the embodied AI benchmarking ecosystem.
|
|
Comment:
|
-
|
|
Reference(s):
|
|
|
Historic references:
|
|
Contact(s):
|
|
| ITU-T A.5 justification(s): |
|
|
|
|
First registration in the WP:
2025-09-09 13:52:43
|
|
Last update:
2026-09-04 14:11:31
|
|