Register and share your invite link to earn from video plays and referrals.

ClaraChengGo
@ClaraChengGo
Babysitting robots with model-ready data at @openmind_agi @fabricFND
4.4K Following    6.8K Followers
Robotics this week is like GPT-6 Astra VS Jev 😂
Who are the most cracked marketing and community talents in SF? Helping frens for hiring.
99% of us still underestimate how GPT-6 Astra changes robotics as an industry.
lmao my fren is asking me to launch a daily streaming session of GRWM: Get Reading (paper) With Me Lowkey wanna try. Anyone interested? 😅
Some notes on loco-manipulation and its data recipe. Helix 2.5 is the first launch of Loco-manipulation these days. Not only uppper-body. So the question is not whether the robot can move and grasp, it is whether it can autonomously decide how to move to better conduct manipulation tasks. Loco-manipulation is hard because the WBC movement is based on the intent of mani. The policy for locomotion should be integrated with mani, then deal with more action space / DoF. The WBC and mani policies will affect each other as well. Imagine you are pushing the drawer when the force on your hand affects your balance or you lean forward to place the force. This process is dynamics so active perception is required. Personally I am curious about the data recipe of Helix 2.5. The blog highlighted that it's pretrained on Index, which looks like stereo data only. But it didn't mention how is the WBC supervision is built. Some previous work covers this as well, like EgoWholeMocap and ZeroWBC. GenRobot published a blog about the pipeline from ego to WBC (the picture below). It adopts vision-only ground-truth capture system and synchronizes an Ego headset with third-person cameras observing the same behavior.
Show more
Today we’re releasing Helix 2.5 We rented 30 homes in the Bay Area. The robots arrived with no additional training and started doing useful work
Noteworthy for long-horizon tasks: >> Completion Signal: when move to next sub-task. Binary / process score / LLM to verify >> Failure Recovery: like @SkildAI highlights tracking progress + recovering from mistakes >> Key frames: whether those are precisely annotated and its tokens carry more weight >> Object-state transition: episodes in a graph
Show more
Quick note on data annotation for robotics ICL. For manipulation model we construct 3D pose to give clear action demonstration. Similarly, I guess we need clear intent demonstration as well, esp for long-horizon task and VLM to call action expert. In @SkildAI S1 blog, it says "since the demonstration may come from a different scene, viewpoint, or embodiment, the policy must implicitly learn the demonstrator's intent, functional correspondences, and task progress to predict the appropriate actions". Basically it introduces some variation and let the model to learn the core intent shared by those variation. Here are 3 types of annotation that helps to specify the intent: 1. outcome annotation: no matter the procedure, the outcome is stable. e.g. ON(mug,table)→INSIDE(mug,dishwasher) 2. role annotation: like the demo in S1 blog, you tell the robot to water the plant with a can but there is only a cup, then the robot knows to use the cup bc they share the same functional role (as the picture shows) 3. trajectory variation under same intent: teach them that intent is not the trajectory. Carry out a list of episodes with same goal but different trajectory
Show more
I brought top researchers back to my Alma mater and hosted an amazing robotics event! throughout the event there were so many amazing ideas that I kept passing the mics like this:
Show more
Last Friday in Hong Kong we ran a closed-room working session across the robotics stack: foundation models, data engines, hardware, and on-site deployment. Researchers and builders flew in from multiple cities for this room specifically. The academic density was high, and the discussion stayed technical: distribution shift, contact-rich control, data pipelines, and the gap between a lab prototype and a system that survives outside campus. Here is the recap.
Show more
Great opportunities! Don't miss out.
#NVIDIA# PhD Research Internship Opportunities for 2027 My team at @NVIDIARobotics is hiring PhD Research Interns for 2027, with possible locations in Santa Clara (USA) and Shanghai (China). Drop me a line if you're interested in hybrid simulation, multi-modal perception, agentic framework, world action models, RL post-training, sim2real research, or any relevant fields.
Show more
I am indeed impressed by @RewardAI_ , huge congrats to @zipengfu @chenwang_j @yifengzhu_ut ! Everyone knows the demo is super impressive that multiple embodiment are working autonomously with excellent speed. But fundamentally it's the omni-embodiment model with model-level abstraction that supports manipulation-level precision. Cross-embodiment is not new. But previously we use retargeting. We first translate human actions into embodiment-specific robot actions before training the policy. But now Reward AI is doing this at the policy level. The representation is not published yet but I expect it's the same policy learns an embodiment-agnostic action representation from human data, and that representation can then be executed across different robot embodiments through their respective control layers. Purely-human data is not new. For example, @GeneralistAI GEN-1 pretrain stage takes 0 robot data. @sundayrobotics takes the similar approach for it's ACT-1 (but added robot data in ACT-2). But it's new that Reward AI achieves cross-embodiment without embodiment-specific retargeting, and with that speed? I am indeed curious. So I think here are the noteworthy in its upcoming blog: Q1: how is the action representation in this "single policy"? How is that cross-embodiment representation different from the traditional one? Q2: how does human data enter the pipeline? How is the synchronization between multimodal data with different Hz? Q3: Please notice that although it aims to achieve the abstraction on data- and policy-level, it didn't kill the embodiment-specific adaptation, but push it to the action layer. Then how heavy is the work to integrate each embodiment? Although it depends on my Q1, everyone who has done that knows it's not a light work.
Show more
Introducing OM-1, our first robot foundation model, zero-shot generalizing to any robot: table-top arms, industrial arms and humanoids. - learned directly from human manipulation data - no teleop/robot data - close to human-level dexterity and efficiency - multi-robot collab
Show more
Great work that measures GPT6 Astra as an Embodied Policy by Yu-Mool Shu and Lipxin Zheng. It further proves with numbers that VLA handles dexterous movement, while the General Model oversees supervision, error correction, and replanning. Interesting numbers: For the average performance across 10 tasks in RoboDojo, PI0.5 gets 24%, Astra gets 37% but together they achieves 62%. But for RoboLab, Astra itself gets 98%. As for the contribution, GPT 6 Astra is correcting only 14.4% of the executed control steps and the remaining 85.6% following the actions by PI0.5​. Link:
Show more
The update of Anthropic when it's pacing the frontier: > Secured up to 5 GW of AI compute, with $100B+ committed to AWS technologies over 10 years. >Secured ~5 GW of next-generation TPU capacity with Broadcom and Google. > 300+ MW of compute powered by 220,000+ NVIDIA GPUs with SpaceX > $30B of Microsoft Azure compute capacity with Nvidia Just some number.
Show more
Take a breath and let's see how exactly other evolutionary tech "pace their frontier". It's not only for AI. The growth for an evolutionary tech usually comes its creators turned to another side (at least it seems so) and advocate to slow down. Nuclear controls the supply chain. We realised in Hiroshima, Nagasaki and Cuban, then paced its fronteir by controlling uranium and plutonium. Gene editing takes tier. We realised when we see the gene-edited baby IRL and paced it frontier by continuing omatic gene editing but not for heritable germline editing. It's not the first time for us to stand at the edge of the frontier and hope to pace something ambiguous but mind-blowing and transformative. I am not Dario, Sam or Elon. But I predict similar patterns happen when we build up: (1) classification: discuss the regulatory case by case (2) permission: liscense / commitee, this is where new authority will form (3) monitor: publicity and transparency But one ultimate difference is that AI swallows the whole stack. A more powerful nuclear bomb can't enable people to produce more uranium. A gene-edited baby can't further edit the genes by itself. But a new version of model swallows what's been built across the full stack, recursively. For the first time we have something with RSI. That's what indeed matters.
Show more
Don’t miss out this piece if you are sending your robots to work
I wanted to find whether GPT-6 hits its limit on robot dexterity. So I just gave it two videos, without states or actions, and told it to do real2sim and physical retarget to Wuji hands. Here are the successful physical rollouts — a Rubik's cube, then an egocentric human video.
Show more
Even more solid looking back from today
This is what robotics policy training is heading to. How it works now: engineer - reward function - mass training - success But Nvidia ASPIRE is more like: Receive a new task - Execute the policy - Evaluate the outcome - Identify failure patterns - Optimize the policy - Repeat - Master the task Self-improvement for robotics model.
Show more
Glad to invite Dr.Steve Cousins from Stanford Robotics Center and host the panel back in HKU, my alma mater! See you guys tmrw!
Panel 1: Building General-Purpose Robot Intelligence: Data, Models, World Understanding & Generalization - Yuyang Cai, Robotics Technical Product Manager at BrainCo - Su Yang, Co-founder & Chief Al Architect at Linkerbot - Xiangyu Chen, 2025/2026 ICRA Competition champion - Yoran Beisher, Project Lead at Mecka AI - Clara Cheng, Core Contributor at OpenMind & Fabric Foundation
Show more
It's official! OpenMind is hosting a Smart Fleet Meet & Greet with Bay Area Robotics Association on October 6th 5PM - 8PM during SF Tech Week @Techweek_ Come meet industry leaders in physical AI and our smart fleet powered by OM1. RSVP below via Partiful.
Show more
The deployment of robotics data collection starts from Tier 1 cities in US like @joinshift @taurobots but from Tier 3/4 cities in China. Guess why.
I tested out the egocentric headset with wrist camera + UMI gripper from @Maniformer. Maniformer is the robotics data company incubated by @AGIBOTofficial . Now it has 2 product lines for data collection: egocentric headset with wrist cameras + UMI gripper. The MEgo View comes with a 5-camera headset + 2 wrist cameras, 300 degree view, global shutter, 2-hour battery life with hot swapping. The MEgo Gripper comes with 3 cameras, optional 3D tactile, global shutter, 200 degree fisheye. For both devices, you can upload the data by Wifi, USB or SD card. The data output includes 7-camera RGB image, depth info, IMU, trajectory and audio. On top of these, the gripper provides tactile info, gripper distance and angle as well. The app expands multiple use cases. It is not only for enterprises to sync the equipments and check the video stream in real time, but also for consumers to claim the tasks and earn earn real cash. Check the video for details. DM is open.
Show more