注册并分享邀请链接,可获得视频播放与邀请奖励。

Zhihu Frontier
@ZhihuFrontier
🚀Bringing China's AI & tech trends, voices and perspectives to the global stage. ⚡️Powered by 知乎/ China's leading knowledge community.
加入 June 2025
180 正在关注    11.6K 粉丝
🦾 Why Data, Not Models, Is the Real Moat in Embodied AI The timing of this question is hard to miss. This week, the World Humanoid Robot Games released a 2,500+ hour dataset covering 12 scenario categories, 44 operations, and more than 10,000 tasks. Crucially, it also includes failures and edge cases. But raw hours tell only part of the story. Zhihu contributor 于超, an assistant professor at Tsinghua Shenzhen International Graduate School, shares his team’s view on why data has become embodied AI’s hardest-to-replicate advantage. 1️⃣ Robot scaling is fundamentally asymmetric Like language models, robot policies appear to benefit from more data, larger models, and greater compute. But these three inputs do not scale equally. Model architectures can be studied and reproduced quickly. General-purpose compute can, in principle, be purchased. High-quality robot data is different. It must be accumulated through physical interaction, and competitors cannot recreate it overnight. That asymmetry is what turns data into a moat. 2️⃣ Robot data must be manufactured LLMs inherited decades of internet data. Robots did not. Every useful trajectory must be produced through physical interaction. Even a simple cup-grasping task changes with the object, lighting, environment, camera angle, and robot body. So raw hours are not enough. What matters is the diversity of embodiments, tasks, objects, failures, and recoveries. Open X-Embodiment needed more than 20 institutions and 22 robot platforms to collect over one million trajectories. DROID used 50 collectors for a year, producing only 350 hours of data. New methods such as UMI and egocentric recording make collection easier. But every hour still requires real people, equipment, and time. 3️⃣ A successful trajectory can still be bad data Robot data quality is more complicated than whether a task was completed. On the hardware side, camera accuracy, encoder readings, force sensors, calibration, communication latency, and synchronization across modalities can all corrupt a trajectory. The human operator adds another source of noise. Teleoperating a robot is not the same as performing the action directly. Operators hesitate, pause, readjust, and develop habits for compensating for the control system. A task may succeed even when parts of the demonstration should never be imitated. Success is therefore only the coarsest possible label. One trajectory can contain both excellent behavior and inefficient or misleading actions. Training on the entire trajectory without distinction effectively tells the robot to learn both. 4️⃣ The next challenge is information density Collecting more trajectories is only half the problem. Teams must also identify which parts are worth learning from. Yu Chao’s team developed STEAM to detect local progress within a trajectory without frame-by-frame annotation or manually designed rewards. It separates useful progress from hesitation, failure, and recovery. The key question is shifting from “How many trajectories do we have?” to “How much useful information does each trajectory contain?” 5️⃣ Embodied data is physically expensive Text can be copied. Videos can be downloaded. Robot data requires a physical production process. Collecting one hour may involve a robot, sensors, teleoperation equipment, an operator, a suitable environment, task materials, and engineers who maintain and calibrate the system. Real factories, stores, and homes add even more complexity. And pressing the record button is only the beginning. Transmission, cleaning, governance, and storage can cost more than collection itself. The author offers a rough calculation. If a company wants one million hours of real-world data and reduces the combined collection and management cost to RMB 200 per hour, the total still reaches RMB 200 million. 🔑 The real moat compounds over time Quantity, quality, and cost explain why embodied AI data cannot be replicated through a short burst of spending. Large, diverse, high-quality datasets require physical infrastructure, operational discipline, and years of accumulation. As robot policies continue to benefit from scaling, the durable advantage will belong to teams that can repeatedly: 🔹 Collect broader real-world experience 🔹 Identify the most informative behavior 🔹 Preserve failures and recovery signals 🔹 Turn noisy trajectories into useful learning data In embodied AI, having data and knowing how to use it are becoming two very different capabilities. And the second may ultimately matter even more than the first. 🔗 Full analysis: #EmbodiedAI# #Robotics# #PhysicalAI# #RobotLearning# #AIData# #ScalingLaw# #Tsinghua#
显示更多