Register and share your invite link to earn from video plays and referrals.

Search results for RobotLearning
RobotLearning community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including RobotLearning
Built on our Data Foundation Model (DFM), delivering higher-quality, more scalable, and more efficient Ego+Fingers data production.Industry-first DFM enables end-to-end data processing#EmbodiedAI# #HumanData# #6DPose# #PhysicalAI# #RobotLearning# #EgoData#
Show more
🤖 Great to meet robotics builders at our Embodied Intelligence Workshop at RoseLab, Toulouse! With SenseCraft Robotics, you can bring the robot arm workflow into one platform:🔗 Connect → 📊 Collect Data → 🧠 Train → 🚀 Deploy 🎁 New users get 1,000 free credits on first login. 👉 #SenseCraftRobotics# #SeeedStudio# #Robotics# #RobotLearning# #EmbodiedAI#
Show more
🤖 Expensive robot data is scarce, so why not learn from our everyday first-person videos? This work turns human hand motion into robot actions to pretrain VLA models. 📰 Title: ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining 🔗 URL: 💡 Overview ACE-Ego-0 unifies robot demonstrations with egocentric human videos (Ego4D, EPIC-KITCHENS, and more) to pretrain Vision-Language-Action (VLA) models, trained on over 6,000 hours of combined data. 🔍 Challenges Solved Robot demonstrations are costly to collect, while human videos are cheap and abundant. But the two differ in action space, embodiment structure, temporal dynamics, and supervision quality, so naively mixing them breaks training. 🛠 Methodology & Proposed Approach ・Unifies actions in head-camera coordinates with 6D rotations, treating the human hand as an end-effector ・Encodes robot URDFs into morphology tokens via a GNN to absorb structural differences ・Chunks actions by consistent physical duration instead of fixed steps for temporal alignment ・Applies a reliability-weighted loss to noisy human videos, focusing on trustworthy position channels 📊 Use Cases / Results On RoboCasa it hits 72.8% average success (vs GR00T-N1.6 at 47.6%) and ~91% on RoboTwin 2.0. On a real bimanual robot it reaches 78.3% average (π0.5: 71.7%). Strikingly, on a task with only 34 robot demos, adding 419 human video episodes lifted success from 10% to 40%, a 4x gain. #RobotLearning# #VLA#
Show more
🦾 Why Data, Not Models, Is the Real Moat in Embodied AI The timing of this question is hard to miss. This week, the World Humanoid Robot Games released a 2,500+ hour dataset covering 12 scenario categories, 44 operations, and more than 10,000 tasks. Crucially, it also includes failures and edge cases. But raw hours tell only part of the story. Zhihu contributor 于超, an assistant professor at Tsinghua Shenzhen International Graduate School, shares his team’s view on why data has become embodied AI’s hardest-to-replicate advantage. 1️⃣ Robot scaling is fundamentally asymmetric Like language models, robot policies appear to benefit from more data, larger models, and greater compute. But these three inputs do not scale equally. Model architectures can be studied and reproduced quickly. General-purpose compute can, in principle, be purchased. High-quality robot data is different. It must be accumulated through physical interaction, and competitors cannot recreate it overnight. That asymmetry is what turns data into a moat. 2️⃣ Robot data must be manufactured LLMs inherited decades of internet data. Robots did not. Every useful trajectory must be produced through physical interaction. Even a simple cup-grasping task changes with the object, lighting, environment, camera angle, and robot body. So raw hours are not enough. What matters is the diversity of embodiments, tasks, objects, failures, and recoveries. Open X-Embodiment needed more than 20 institutions and 22 robot platforms to collect over one million trajectories. DROID used 50 collectors for a year, producing only 350 hours of data. New methods such as UMI and egocentric recording make collection easier. But every hour still requires real people, equipment, and time. 3️⃣ A successful trajectory can still be bad data Robot data quality is more complicated than whether a task was completed. On the hardware side, camera accuracy, encoder readings, force sensors, calibration, communication latency, and synchronization across modalities can all corrupt a trajectory. The human operator adds another source of noise. Teleoperating a robot is not the same as performing the action directly. Operators hesitate, pause, readjust, and develop habits for compensating for the control system. A task may succeed even when parts of the demonstration should never be imitated. Success is therefore only the coarsest possible label. One trajectory can contain both excellent behavior and inefficient or misleading actions. Training on the entire trajectory without distinction effectively tells the robot to learn both. 4️⃣ The next challenge is information density Collecting more trajectories is only half the problem. Teams must also identify which parts are worth learning from. Yu Chao’s team developed STEAM to detect local progress within a trajectory without frame-by-frame annotation or manually designed rewards. It separates useful progress from hesitation, failure, and recovery. The key question is shifting from “How many trajectories do we have?” to “How much useful information does each trajectory contain?” 5️⃣ Embodied data is physically expensive Text can be copied. Videos can be downloaded. Robot data requires a physical production process. Collecting one hour may involve a robot, sensors, teleoperation equipment, an operator, a suitable environment, task materials, and engineers who maintain and calibrate the system. Real factories, stores, and homes add even more complexity. And pressing the record button is only the beginning. Transmission, cleaning, governance, and storage can cost more than collection itself. The author offers a rough calculation. If a company wants one million hours of real-world data and reduces the combined collection and management cost to RMB 200 per hour, the total still reaches RMB 200 million. 🔑 The real moat compounds over time Quantity, quality, and cost explain why embodied AI data cannot be replicated through a short burst of spending. Large, diverse, high-quality datasets require physical infrastructure, operational discipline, and years of accumulation. As robot policies continue to benefit from scaling, the durable advantage will belong to teams that can repeatedly: 🔹 Collect broader real-world experience 🔹 Identify the most informative behavior 🔹 Preserve failures and recovery signals 🔹 Turn noisy trajectories into useful learning data In embodied AI, having data and knowing how to use it are becoming two very different capabilities. And the second may ultimately matter even more than the first. 🔗 Full analysis: #EmbodiedAI# #Robotics# #PhysicalAI# #RobotLearning# #AIData# #ScalingLaw# #Tsinghua#
Show more
This robot is skateboarding completely blind. And yes... It’s autonomous. No vision or markers. Just proprioception, joint torques and some very clever state estimation. Aditya Bhatt (PhD at IAS / DFKI under Jan Peters) just dropped one of the wildest robot demos of the year. People in the comments are already losing it over the same questions: How the hell does it know where the deck is? Pure IMU? Contact models? What observations is the policy actually getting at inference time? They only trigger the tricks like a real-life video game. This is coming straight out of @Jan_R_Peters’ lab, @ias_tudarmstadt. The same Jan Peters who just sat down for a long, unusually honest interview about how German red tape cost the field a full five-year head start on humanoids. If you’re into real robot learning (not just pretty videos), both are worth your time. Skateboarder demo and work by @aditya_bhatt. Podcast: Ep 105 | Red Tape Cost Us a 5-Year Lead in Humanoid Robotics (w/ Jan Peters) Huge thanks to Jan Peters for the rare honesty, his humour and insights into the world of robotics and his own! Full Conversation: On Spotify: On YouTube:
Show more
“In-Context Robot Learning with VLM Agents” Most robots need lots of task-specific training before they can do something new. This paper shows you can instead just show a general VLM what to do and let it adapt directly from context. GPT-Policy essentially turns robot learning into prompting. So they gave the frozen VLM an example of the task, even just a human video with no robot action labels, and it can translate what it sees into robot actions on the fly. This brings few-shot learning into the physical world, where teaching a robot a new behavior could look more like showing it an example than retraining a policy.
Show more
WAMs and VLAs for Robot Learning | Cosmos Labs
WAMs and VLAs for Robot Learning | Cosmos Labs
From Digital Twins to Data Engine: Cutting the Real-Data Burden with Sim-Powered Robot Learning Teaching a robot a new task could take hundreds of teleoperated demonstrations. For foundation models, adapting to an entirely new robot can cost orders of magnitude more — dedicated hardware, trained operators, months of engineering. On @boosterobotics' dual-arm robot, we studied this at two levels: ✱ Specialist: Can task-aligned simulation mixed with a small set of real demonstrations reduce the real-data burden? ✅ Yes. With only 10 real demos, the policy made no contact at all in physical rollouts (0/20). Adding 50 simulated trajectories brought contact to 17/20. ✱ Foundation: Can data accumulated across tasks build a reusable starting point (a Booster-specific model prior)? ✅ Yes. After full-parameter continued pretraining, a model adapted with just 30 demonstrations per task beat the original given twice as many: 14/16 vs 10/16 in simulated evaluation. Before any task-specific adaptation, in zero-shot simulation, it was already roughly 3× closer to the target (17.27 cm → 5.78 cm). This work runs on Axis Suite, our Physical AI solution across different robot embodiments. Distributed contributors generate task-aligned sim data on Axis Hub at scale, reducing real-data needs for specialist adaptation while powering cross-embodiment generalist training. Read the full blog:
Show more
0
129
445
61
Forward to community
Reward AI is based in the SF Bay Area. The founders are two of Stanford's standout robot-learning researchers: Zipeng Fu (Cofounder, CEO), advised by Chelsea Finn, ex-Google DeepMind. Lead author on Mobile ALOHA (the viral $32k robot that cooks and cleans) and HumanPlus (humanoids learning by shadowing humans). Chen Wang (Cofounder, CTO), advised by Fei-Fei Li and Karen Liu. Lead author on DexCap, the portable dexterous-manipulation capture system OM-1's wearable hand is built on.
Show more