China's National Data Administration moves to set standards for the data that trains robots
China's National Data Administration, the state body that runs the country's data policy, said after its head Liu Liehong met robotics companies and research institutes on 10 September that it would push standards for embodied-AI data in due course, guide local data authorities and support companies' investment in data. Ten days earlier seven companies had asked the same agency for public data infrastructure and common standards; by one Chinese institute's estimate, robot foundation models need roughly 10 million hours of real-world training data, and far less exists.
Artificial Intelligence··Morning
Liu Liehong's 10 September meeting
China's National Data Administration, the state body that runs the country's data policy, described in a statement posted on its own site on 13 September a meeting chaired by its head Liu Liehong on 10 September. Representatives of the Chinese Academy of Sciences' Institute of Automation, the Beijing Institute for General Artificial Intelligence, Lightwheel, JD Group, Galaxea, Spirit AI, the Shijingshan humanoid-robot data training centre and the Beijing Humanoid Robot Innovation Center attended and gave their views under the heading of data empowering embodied AI. Embodied AI refers to systems such as robots that perceive and act through a physical body. Deputy head Xia Bing also attended. In The Next Web's account, research institutes, technology companies and humanoid robotics organisations took part, and the agency said the industry is becoming more data-driven.[1], [2]
The administration's pledge: standards, guidance for local authorities and investment support
In the statement the administration said the industry is data-driven and that demand for high-quality, diverse and large-scale data keeps growing. It said it would push standards for embodied-AI data in due course, work problem by problem, strengthen planning and guide local data authorities to carry out the work in an orderly way. The Next Web likewise reported that the agency would write standards for the data that trains robots, guide local authorities on the work and support companies putting more money into data resources.[1], [2]
Ten days earlier seven companies had asked for exactly this
The industry asked first, in The Next Web's account: ten days earlier the same agency had met seven companies, which requested public data infrastructure for embodied AI and common data standards. The size of the data gap appears in one estimate: by the reckoning of the China Academy of Information and Communications Technology, embodied-AI foundation models need roughly 10 million hours of real training data, while between 100,000 and 1 million hours of high-quality data exist worldwide.[2], [1]