Register and share your invite link to earn from video plays and referrals.

Search results for EgocentricAI
EgocentricAI community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including EgocentricAI
🧵 New repo: Awesome Egocentric Dataset — a curated list of first-person vision datasets, refreshed beyond the classic Ego4D-era list. ✅ Recent 2023–2026 releases (Ego-Exo4D, EgoSchema, AEA, EgoExoLearn, EgoLife…) ✅ Dead/broken links removed, migrated ones fixed ✅ A section on egocentric-driven VLA papers (training source + data-processing), with first-principles + fishbone analysis 🚀 Open-source, contributions welcome. 🔗 #EgocentricVision# #FirstPersonVision# #VLA# #ComputerVision# #Datasets#
Show more
the most famous egocentric datasets are people cooking in their own kitchens. the robots we're training on them are headed for warehouses, garages, and factory floors apac egocentric stereo is 12 first-person recordings of people actually doing their jobs: an automotive garage, a construction site, an electronics factory, a bar, a shipment hub, a laundromat head-mounted stereo rig, 1920x1080 per eye at 30 fps, plus a depth render, hand and head tracking, and a caption for what the wearer is doing at every moment. 248 segments spanning 62 distinct verbs checkout the dataset parsed into fiftyone format. every stream scrubs on one shared timeline in fiftyone: both eyes, depth, tracking, and captions together, one line to load checkout the dataset here: or just jump right in with the hugging face space hosting the dataset:
Show more
Humanoid AI models are hungry for diverse egocentric data with high-quality hand pose annotations. Ropedia's HOMIE Gen2 targets that demand. A 380g headset that captures 360° from four cameras plus four-channel spatial audio, and outputs depth maps, camera poses, 21 hand and 52 body keypoints, 3D scene reconstruction, and task annotations, all synced to 50 microseconds. High precision: Hand keypoint accuracy at 4.68mm. Action labels are auto-generated and 96% accurate. 13 hours of continuous capture, no external rig. @ropedia_ai
Show more
We took our robotics hardware to @SolanaEvents AI Accelerate Day. Live teleoperation and egocentric data were recorded — with good coffee, Rubik's cube fun, and a community that showed up. 👇
Show more
Another day, another Robotics Lab delivering millions of hours of egocentric human experience and data. Do you remember a couple years ago when everyone was complaining that we didn’t have enough data for robotics? Well, now it’s definitely getting to the point where that bottleneck is starting to disappear. This is why I’m glad there are so many NeoLabs popping up measuring body motion, object interaction, tactile force, 3D structure, etc. (and in Maxinsights’ case, capturing 3D structure instead of just raw first-person video). Their scaling thesis is also pretty interesting an hour of someone repeatedly doing one clean task is not worth the same as an hour full of different objects, contacts, failures, recoveries, tool use, and corrections. So they basically think of it as effective experience = hours × information per hour, meaning the fuck ups and how humans recover from them are part of the valuable data too.
Show more
Most companies that shut down give up their impact. This startup did something much better: they open-sourced 1,274 hours of egocentric robotics data. 13,451 recordings of humans doing everyday tasks, as a gift to the robotics community. Thank you @eidon_ai! Open source has a superpower: work can outlive the organization that created it. More startups should do this!
Show more
0
52
1.4K
141
Forward to community
Lumos Robotics is heading to Pittsburgh for IROS 2026. We’re bringing: Lumos NIX — AI mini bipedal humanoid Lumos NexCore — Skill Evolution Engine Lumos Station Lite — Dual-Arm Visual Manipulation and Embodied AI Development Platform Lumos Ego Vue — Egocentric data collection device 📍 Sep. 28–30 | Booth #536# Meet us at IROS 2026 in Pittsburgh. #IROS2026# #EmbodiedAI# #Robotics# #IntelligenceForTheRealWorld#
Show more
Excited to see LingBot-VLA 2.0 open-sourced by @Robbyant. We're proud to have contributed the human data behind it — egocentric video paired with <1 cm hand tracking captured in the wild at scale and processed through our Data Foundation Model. #EmbodiedAI# #Robotics#
Show more
🤖 Expensive robot data is scarce, so why not learn from our everyday first-person videos? This work turns human hand motion into robot actions to pretrain VLA models. 📰 Title: ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining 🔗 URL: 💡 Overview ACE-Ego-0 unifies robot demonstrations with egocentric human videos (Ego4D, EPIC-KITCHENS, and more) to pretrain Vision-Language-Action (VLA) models, trained on over 6,000 hours of combined data. 🔍 Challenges Solved Robot demonstrations are costly to collect, while human videos are cheap and abundant. But the two differ in action space, embodiment structure, temporal dynamics, and supervision quality, so naively mixing them breaks training. 🛠 Methodology & Proposed Approach ・Unifies actions in head-camera coordinates with 6D rotations, treating the human hand as an end-effector ・Encodes robot URDFs into morphology tokens via a GNN to absorb structural differences ・Chunks actions by consistent physical duration instead of fixed steps for temporal alignment ・Applies a reliability-weighted loss to noisy human videos, focusing on trustworthy position channels 📊 Use Cases / Results On RoboCasa it hits 72.8% average success (vs GR00T-N1.6 at 47.6%) and ~91% on RoboTwin 2.0. On a real bimanual robot it reaches 78.3% average (π0.5: 71.7%). Strikingly, on a task with only 34 robot demos, adding 419 human video episodes lifted success from 10% to 40%, a 4x gain. #RobotLearning# #VLA#
Show more
🕶️ Walk through a first-person world with your own body motion, and explicitly specify what exists at a given location with an image and pose, including how it evolves over time. Meet AnchorWorld, an embodied egocentric world model. Title: AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization URL: 📝 Overview AnchorWorld generates first-person video controlled by full-body human motion. With "anchor views," it lets you explicitly specify what exists at a given 3D location and how it changes over time. ❓ Challenges Solved Existing world models struggle to supervise full-body motion from egocentric video alone, and define environments only implicitly. They lacked both natural embodied control and localized world customization. 💡 Methodology & Proposed Approach ・Since most of the body is invisible in first person, it uses third-person video as auxiliary supervision to learn body-environment positioning ・An anchor has three parts: an RGB image, a 6-DoF viewpoint pose, and an evolution prompt that specify local appearance and temporal change ・3D RoPE spatially distinguishes multiple anchors, and masked cross-attention enables anchor-specific text control ・It trains in four stages (third-person, first-person, static anchors, dynamic evolution), built on Wan 2.2 TI2V 5B 🎯 Use Cases It applies to embodied VR apps, first-person game environment design, embodied-AI training scenarios, and interactive video generation with localized control. 📊 Experimental Results ・On egocentric static scenes it reaches CLIP-V 0.885 and camera accuracy ATE 0.112m, beating PlayerOne and others ・On egocentric dynamic scenes, text alignment (VideoAlign-TA) is 0.717, far above CaM-Ego's 0.385 ・It generalizes strongly to out-of-distribution UE and real-world scenes with little visual overlap between the initial view and anchors #WorldModel# #EmbodiedAI#
Show more