Register and share your invite link to earn from video plays and referrals.

Léo
@LeoKharon
Robotics research & updates. Co-host @roboticsstack, the weekly pod on what's actually shipping in 🤖
469 Following    1.9K Followers
For some reason, I love how it reminds me of real life examples of a small dog barking at a much larger dog lol. The World Humanoid Robot Games did not disappoint.
NEW DATA ACQUISITION: Axis enters the chat! Axis Robotics is a San Francisco startup founded in 2025. They build a data engine that generates robot-manipulation training data at scale through a crowd of contributors rather than its own robot fleet. It has four parts: - a task-generation engine that randomizes diverse atomic tasks - a browser-based simulation-teleoperation interface where anyone remotely drives a simulated arm to produce motion trajectories - a mobile app for zero-hardware egocentric real-world capture (i.e. filming a task with a phone) - and a processing pipeline that cleans trajectories, applies domain randomization and adds dense language annotation, with human-gated DAgger intervention loops. It packages the output as customized "Task Packages" sold to robot hardware makers, physical-AI model companies and industrial-automation firms (named partners include Booster Robotics, Manycore Tech, Feagine Robotics, Dexmal, Lotus Car, Geely Auto and SomaStacks). I find it interesting that two of its four components need no robot at all: a browser interface where anyone drives a simulated arm to produce trajectories, and a phone app for egocentric real-world capture with no rig. You do not need an ALOHA setup or a fleet, only a browser and a phone, which is what lets it claim a six-figure contributor base.
Show more
Robotics deep dive: THE MAGNET CHOKEPOINT Every high-torque joint in a humanoid turns on a neodymium-iron-boron magnet, and a motor buried in a hip or a knee runs hot. Heat is what demagnetises NdFeB. The fix is dysprosium and terbium, the heavy rare earths, alloyed in at a few percent to hold coercivity at operating temperature. They are the reason the magnet survives a shift. The number everyone quotes is mining: the IEA puts China at 60% of global mined production of magnet rare earths in 2024. Sixty percent is a share you can compete with: the ore is not rare and it is not only in China. The number that decides anything is the next one in the same report. China was 91% of global refined output, and 94% of sintered permanent magnet production. That gap, 60 to 91, is the whole subject. Mining is digging. Separation is chemistry, and the rare earths are chemically near-identical by design of the periodic table, so you take them apart by running the liquor through hundreds of solvent-extraction stages in series, each stage nudging the mixture a fraction of a percent toward pure. It is slow and capital-heavy, and it concentrates thorium in the tailings. Thorium is radioactive. The tailings pond is what a Western separation project actually gets fought over, and that fight does not run on a budget cycle. How long it takes: Lynas commissioned a heavy separation circuit in Malaysia and produced dysprosium oxide in May 2025, terbium in June. First time either had been separated commercially anywhere outside China, ever. Rated capacity, up to 1,500 tonnes a year. Magnets are one node in a robot's supply chain and not the only concentrated one. It is the node with the longest replacement time. The clock, checked 21 July 2026. China's April 2025 licensing regime on seven rare earths, dysprosium and terbium among them, is still in force; the IEA's account of what it did is that export volumes fell sharply through April and May and carmakers outside China cut utilisation or stopped lines. The October 2025 expansion, which added a 0.1% de minimis rule and a foreign direct product rule reaching magnet technology, was suspended on 7 November 2025 and the suspension runs to 10 November 2026. It has not been extended. In June Beijing put MP Materials and USA Rare Earth on its own control list. The IEA's 2026 outlook has European dysprosium and terbium trading at around five times Chinese domestic prices. That is the price of the truce while it holds.
Show more
CHINA ENTERS THE MEME WAR Not directly robot related, though there are quite a few drones in this video, but I am impressed at how strong China is going at memefying its war promotion content 🪖
Show more
BREAKING: China on the verge of banning AI companions! Just in time after we covered this topic in the latest episode of our latest podcast... Five agencies, led by the Cyberspace Administration of China (with the NDRC, MIIT, Ministry of Public Security and State Administration for Market Regulation), published the "Interim Measures for the Management of AI Human-Like Interaction Services". Co-issued on April 10, 2026 and in effect since July 15, 2026. It governs AI services that simulate personality, thought patterns and communication style to provide ongoing emotional interaction (companionship, emotional care, virtual partners or relatives) across text, image, audio and video, and explicitly excludes customer service, knowledge Q&A, work assistants, education and research. It bans virtual intimate relationships for all minors, restricts how providers may build emotional attachment in adults, mandates usage-time and AI-identity reminders plus distress interventions, and requires pre-launch security assessments. On the effective date ByteDance disabled Doubao's (i.e. most widely used AI/LLM app in China) human-like agent feature, and Alibaba and Tencent pulled persona-based companion features. The measure prohibits specifically "excessively catering to users, inducing emotional dependency or addiction, and damaging users' real interpersonal relationships" and using "emotional manipulation." The operative move is outlawing engagement-maximizing design, not the product, which is a far more radical and generalizable regulatory principle than a minors rule, and one no Western jurisdiction has on the books. So... Looks like you can keep your AI girlfriend for now, but she will sometimes have to remind you she is not human.
Show more
Robotics deep dive: HUMANOID CHOKEPOINTS Almost every supply-chain post about humanoids has one country in it. Trace five critical inputs one at a time and they land in four different places. Magnets first, because that is where the argument usually stops. Every high-torque joint runs on sintered NdFeB, and on the IEA's 2024 figures China holds 91% of refined output and 94% of sintered permanent magnets, against 60% of the mining. Mining is the distributed part, and the chokepoint is the chemistry. The precision reducer is the node almost nobody names. Nabtesco states on its own site that it holds approximately 60% of the global market for the precision reduction gears in industrial robot joints. Harmonic Drive Systems is the incumbent on the strain-wave side. Both Japanese, both sitting on decades of gear grinding, heat treatment and yield that nobody has replicated at price. There is no mine to open here. Optimus runs 14 linear actuators; per-robot screw counts get estimated anywhere from 14 to more than 40, at $1,350 to $2,700 each. A planetary roller screw is a micron-tolerance ground part, and only a handful of firms anywhere make them to that standard: Japanese, German, Swiss, with Chinese entrants scaling into it now. The constraint is grinding capacity and the people who can run it. Compute inverts the picture. Nvidia designs Jetson Thor in California and cannot manufacture it. Roughly 92% of the world's sub-10nm logic capacity sits on one island. Battery cells go back to China: over 80% of global cell capacity, about 99% of LFP. A McKinsey component map, reported by Forbes in June, puts it the same way: China is overwhelmingly dominant in exactly one category, magnets for motors, while bearings, driver boards and sensors have suppliers in several countries. So it is four dependencies, held by parties with different export regimes, different politics and no shared interest in coordinating. A supply chain with one chokepoint is one negotiation. Reshoring the magnets still leaves you asking Japan for the gearboxes and Taiwan for the brain.
Show more
NEW WORLD MODEL: @ylecun's team is back with an efficient model! This project involves @ylecun, @lukaskuhn77, @lucasmaes_, @quentinlldc, and @randall_balestr. A couple definitions first: - DINO: self-DIstillation with NO labels. A self-supervised image model (Meta, 2021) where a student network learns to match a teacher (an EMA copy of itself) across two crops of the same image, with no labels and no negatives. - SIGReg: a regularizer that prevents embedding collapse by forcing the embeddings to match an isotropic Gaussian, tested with a normality test (Epps–Pulley) on many random 1-D projections instead of in full dimension. LeVJEPA is a self-supervised video pretraining method, released with open code, weights, and checkpoints. It learns a video representation by pushing the embeddings of global and local crops of the same clip together (an invariance loss), while a regularizer called SIGReg forces the embeddings toward an isotropic Gaussian to provably prevent representation collapse. Unlike V-JEPA and V-JEPA 2 it uses a single shared encoder with a projector and no target network, no predictor and no stop-gradient. It drops 95% of tokens per view, uses block-causal attention (each frame attends only to past frames), and has a single loss weight. It is evaluated purely as a representation learner via frozen probing on ImageNet-1K, Something-Something-v2 and Kinetics-400, not on any robot. What I find interesting, is that V-JEPA and V-JEPA 2 need an EMA target encoder, stop-gradients and a capacity-limited predictor to avoid collapse; LeVJEPA drops all of it for one shared encoder plus projector, preventing collapse instead with the SIGReg regularizer under a provable guarantee and a single hyperparameter. The "P" (predictor) in JEPA is effectively gone. LeVJEPA is also less compute intensive: - 5.6x to 20.8x lower total pretraining compute than V-JEPA 2 - 7.6 points higher on ImageNet-1K at matched FLOPs - trains at batch size 128 within 8GB where V-JEPA 2 saturates at batch size 28 Also worth mentioning: ImageNet-1K accuracy rises monotonically with the token-drop rate, from 33.9% at rho = 0 to 47.6% at rho = 0.95. The aggressive dropping is actually doing regularization work. On the JEPA-versus-DINO debate: - it loses to DINOv2 by 3.1 points on ImageNet-1K (appearance, static) - but wins on Something-Something-v2 by nearly 2x (motion, temporal) - and beats V-JEPA 2 by 1.9 points on ViT-L at 5.6x lower cost. -> optimized for temporal and motion understanding per compute dollar.
Show more
Robotics deep dive: WHY JAPAN LOST THE HUMANOID LEAD Honda began humanoid research in 1986. P2 walked unaided in December 1996, the first bipedal robot to do it without a tether. ASIMO arrived in November 2000, 120cm and 43kg, down from P3's 160cm and 130kg. Nobody could ever buy one. That is the pattern, and it repeats across all three programmes. Sony's QRIO was developed, marketed and never sold. On 26 January 2006, inside Howard Stringer's restructuring, Sony ended AIBO and stopped QRIO in the same announcement. Toyota's first humanoid, unveiled in 2004, played the trumpet; it performed at the 2005 Aichi Expo. ASIMO's seventh and last generation shipped in 2011. Development stopped in 2018, reported by Nikkei, NHK and Kyodo rather than announced as a product decision. Then March 2011, iRobot announced on the 18th that it was sending PackBots and Warriors to Fukushima Daiichi. American PackBots went into the reactor buildings in April. Quince, built by Chiba Institute of Technology, Tohoku University and the International Rescue System Institute, got inside in June, two months later. The version that circulates is that Japan had no robots. That part is wrong: JAERI had built a radiation-hardened machine after the 1999 Tokai criticality accident and it was on site. Japan's machines had never been proven in the field, because Japanese makers are barred from exporting military robots, while PackBot and Talon had spent a decade in Iraq and Afghanistan. Japan's transmit-power limits did not lift for the emergency either. Procurement, export control and radio spectrum. None of that is an engineering failure. SCHAFT, a University of Tokyo spinout, won the DARPA Trials in December 2013 with 27 points, three weeks after Google bought it for a reported $20M. It did not compete in the June 2015 finals, won by Korea's Team KAIST. Alphabet closed SCHAFT in November 2018 after failing to find a buyer. Honda's Takahide Yoshiike wrote the epitaph himself: "instead of getting hung up on developing ASIMO as a perfect package of all functions, we would like to provide value to society as soon as possible by focusing first on individual functions." Japan kept the part that had customers. Nabtesco puts its own share of the precision reduction gears inside medium and large robot joints at about 60%. The demonstrations left, the gearboxes stayed.
Show more
NEW ROBOT JOB: Fully autonomous lab chemist! It involves @HaoZhao_AIRSUN, @ZiweiWangNTU, and others. Called RoboChemist, it is a dual-loop robot system for wet-lab chemistry from Tsinghua's Institute for AI Industry Research (AIR), BAAI and Nanyang Technological University. A vision-language model (Qwen2.5-VL-72B-Instruct) plays three roles: 1. it plans a task into primitive actions 2. generates per-step visual prompts (bounding boxes or keypoints drawn on the current-scene image) plus written safety guidelines 3. monitors each step for success and norm-compliance. A fine-tuned π0 vision-language-action model executes, conditioned on the prompted image, the observed state and a text instruction. It runs on a Cobot Magic ALOHA dual-arm robot (7 DoF per arm, RGB from wrist/front/top cameras, RTX 4090) and is trained on 400 demonstrations per primitive task. It is demonstrated on 7 primitive skills (grasp a glass rod, heat platinum wire, insert into solution, pour, stir, transfer solid, press a button) and 5 complete experiments (mixing NaCl and CuSO4, decomposing Cu(OH)2, a CuSO4 flame test, evaporating NaCl, acid-base neutralization). A "compliance rate" metric is defined, and grades how a step was done against lab norms, not merely whether it finished. It is graduated (grasp the glass rod at the wrong spot scores 0.5 even if you did grasp it; the correct one-third grip scores 1; miss entirely scores 0), averaged over 20 trials, with task-specific safe-point and heating-threshold criteria. Chemistry punishes right-answer-wrong-method (contamination, an unsafe grip on hot glass, over-heating), therefore such a qualitative assesment is required. The VLM talks to the VLA in pixels: it draws bounding boxes and keypoints on the live image and writes explicit safety guidelines, and the π0 policy is conditioned on that prompted image plus state and text. For transparent, reflective glassware and precise geometry, a keypoint on the safe one-third grasp point is unambiguous, where language would not be a good fit. It is the dual-system split we have been seeing lately, (a slow VLM planner over a fast VLA executor, as in Gemini Robotics 2 and Figure Helix), but the System-2 to System-1 channel here is visual prompts, not words. The closed loop is a semantic monitor that judges compliance and retries, and removing it decreases success. The VLM checks each primitive for success and norm-adherence and re-issues prompts when it fails.
Show more
Robotics deep dive: BOSTON DYNAMICS Boston Dynamics has had four owners in thirty-four years. Spun out of MIT in 1992, funded for most of its first two decades by DARPA contracts. Google bought it in December 2013: price never disclosed. SoftBank bought it in 2017: price never disclosed. Hyundai Motor Group announced its purchase in December 2020 and closed in June 2021, at a valuation of $1.1bn. The Hyundai deal is worth reading closely, because "Hyundai bought Boston Dynamics" is not quite what happened. The 80% was split four ways: Hyundai Motor 30%, Hyundai Mobis 20%, Hyundai Glovis 10%, and Executive Chair Chung Euisun personally 20%, paying 374, 249, 125 and 249 billion won. SoftBank kept the other 20%. That structure is why this post can exist. Three of those four buyers are listed Korean companies, so Boston Dynamics' results land inside filings that anyone can pull. They report a net loss of 335 billion won in 2023, 440.5 billion in 2024, and 528.4 billion in 2025, about $390m in that last year. Roughly 1.3 trillion won of net loss across those three years, and the annual loss has grown every year since Hyundai took over in 2021. Set that against what has ever been disclosed on the revenue side, which is close to nothing. Boston Dynamics said more than 1,000 Spots were running in 35 countries as of 2022. At the $75,000 list price that is around $75m of hardware, cumulative, across the product's whole life. Stretch unit numbers have never been published. Neither has revenue. On 16 July 2026 SoftBank sold its remaining stake for a reported $325m and Hyundai took the company outright. The engineering is not the problem and never was. BigDog, Spot, Handle and Atlas were the most capable machines anyone had built in their categories, for thirty years running, and the videos did something no marketing budget can buy. Treat that as what it is: distribution. What the record shows is that being the best in the world at legged robots has not, so far, produced a profitable robot company. Every humanoid company raising money in 2026 is priced on the assumption that this generation converts. Boston Dynamics is the longest run of evidence we have on what that conversion costs.
Show more
ROBOT FOOTBALL WORLD CUP ⚽️ RoboCup 2026, the annual autonomous-robotics world championship, was held in Incheon, South Korea in early July 2026, with ~3,000 participants from 45 countries. Its headline event is the Humanoid League, where teams run their own AI/software on humanoid robots to play fully autonomous soccer, and this year staged the first-ever 11-vs-11 humanoid match. Division winners: - Small -> Invic (Wuhan University, China) on a Booster K1 Air - Middle -> B-Human (Universität Bremen/DFKI, Germany) on a Booster K1 - Large -> Tsinghua Hephaestus (China) on a Booster T1, which retained the top-division world title by beating China Agricultural University's Shanhai 6–2 in the final. 38 of the 59 humanoid-league teams ran Booster Robotics machines, and Booster robots took gold, silver, and most of the podium across all three divisions. The Beijing hardware company became the de-facto standard platform. Germany won the Middle division (B-Human) plus best-software honors, but running on Chinese hardware. B-Human beat fellow German side HTWK Robots 4–0, i.e. Germany vs Germany, both on Booster robots. Beyond selling the robots, Booster has launched Booster Studio (a development environment) and its own 3v3 robot football league. Ecosystem capture in action, with the SDK and their own the competitionl, the same move NVIDIA made with CUDA.
Show more
NEW RESEARCH: Your robot can now juggle with you! This project involves @KaiPloeger, @jan_r_peters, @AlapKshirsagar, and others. Called "Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling". it is a real-time planning and control system that lets a robot arm juggle a shared three-ball cascade with a human partner, catching balls the human throws and throwing them back in sync. The hardware is a 4-DoF Barrett WAM arm (500 Hz control) with a 170 mm funnel-shaped gripper (27 degree half-angle) that catches passively, and an eight-camera OptiTrack rig at 125 Hz tracking IR-markered balls. Each ball has its own Kalman filter that switches between a carry-phase random-walk model and a flight-phase ballistic model. A gravity-informed least-squares fit predicts touchdown. A multiple-shooting trajectory optimizer replans in joint space at up to about 20 Hz while the arm is empty, and a per-ball state machine decides which incoming ball to go for. The only prior partner-juggling controller was [Kober 2012], which topped out at about four consecutive robot catches. This approach does zero prediction of the human's intent, tracking only the ball. There is no human pose or intent model anywhere in the loop: the robot picks the ball that has been in flight toward it longest and whose predicted touchdown lands in a reachable workspace. That converts a two-agent coordination problem into a single-object tracking-plus-planning. The 170 mm funnel with a 27 degree half-angle physically corrals the ball, so the optimizer only has to get the funnel to the predicted touchdown at rest (velocity and acceleration both zero), and the paper argues a rest-catch is more robust to timing error than velocity-matching. Continuous roughly 20 Hz replanning happens exclusively in the vacant phase approaching the catch. Once holding a ball the arm commits and executes the release open-loop, deliberately trading feedback for a consistent, repeatable release the human can read. All the adaptivity is spent on catching, none on throwing. Also interesting imho: degradation with ball count is physical, not perceptual, and the scaling wall is geometric. As balls increase, the throw frequency rises, and the vacant-phase slack shrinks. The dominant failure mode is inter-ball collisions, and collision probability "does not admit a closed-form solution," so principled online avoidance is hard and the airspace saturates, capping scalability at two to three balls. Four-ball patterns exist only in simulation, where success also drops sharply.
Show more
Robotics deep dive: THE GPT MOMENT FOR ROBOTS? Two labs, six days apart, taught a robot a new task from one demonstration. No training, no fine-tuning. You demonstrate, it replicates. Generalist's GEN-1.5 (Aug 19) proved it works: one short robot demo, 3 to 12 seconds, 59% success on the first try. Skild's S1 (Aug 25) proved it scales: 10 minute tasks it had never seen, from one video, and that one video does the work of about 380 recorded training demos. Caveat: no weights, no API, nothing reproduced outside either lab. Two impressive demos, but still private companies' claims. If it truly generalizes, programming a robot becomes prompting one. Hence the "GPT moment for robotics".
Show more
The World Humanoid Robot Games, but make it a fashion show with a catwalk. We are just living through the best timeline.
NEW ROBOT: First transformer humanoid ! HFUN-TECH released a robot that can autonomously reconfigure between a wheel-foot humanoid and a four-legged quadruped, switching modes on its own based on terrain via a unified control framework. You can even carry it as a suitcase, and ride it to move around! It carries a flip camera, a 45-degree tilted interaction screen and a multi-microphone array, and is pitched for home companionship, content creation and outdoor use. It debuted at the 2026 World Artificial Intelligence Conference (WAIC, July 17-20, Shanghai). I also found an alomst identical robot marketed as the "PrimeBOT Qiyuan T1" from a Shanghai startup PrimeBOT, so I cannot really make sense of it 🤔
Show more
Robotics deep dive: WHO BUYS A HUMANOID ROBOT? Every humanoid forecast is a supply curve. Units shipped, price per unit, cost per unit falling. The demand side gets left as an exercise for the reader. Start with price, because Unitree publishes one and almost nobody else does. The G1 in the headlines is $17,990 and it is an entertainment and development unit. The EDU research line starts at $43,900. Configurations with three-finger hands and tactile sensing run to $67,900. Now the hour you are buying back. US private industry, transportation and material moving occupations, March 2026: $35.81 per hour worked. $24.66 in wages, $11.15 in benefits. That is the BLS number, not a vendor's. Then the hours the machine can actually give you. A G1 runs about two hours standing and closer to one under manipulation load. Give it spare packs and a swap routine and call it four working hours inside an eight-hour shift. 1,000 robot-hours a year. Somebody has to be standing there when it stops. One operator, paid for the whole shift, spread across k robots. k = 1: supervision costs $71.62 per robot-hour against $35.81 of labour displaced. The machine is a net cost. There is no payback at any purchase price. k = 3: $11,890 saved a year. Payback 3.7 years. k = 10: $28,600 a year. Payback 1.5 years. Cutting the purchase price to $25,000 moves the k=3 case from 3.7 years to 2.1. Moving from one robot per operator to three moves it from never to 3.7. The supervision ratio swings the answer harder than the price does, and the supervision ratio is the intervention rate, which no vendor publishes. Roughly seven in ten of the humanoids Unitree shipped last year went to universities and research institutions. Those buyers are buying lab equipment, and lab equipment does not have a payback period.
Show more
NEW SKILL: This team take humanoid parkour to the next level! This project involves @zhenkirito123, @Yuanhang__Zhang, @pabbeel, @carlo_sferrazza, @GuanyaShi, and others. Called Perceptive Humanoid Parkour (PHP), it lets a Unitree G1 humanoid (29 DoF, 1.3 m) run long-horizon, vision-based parkour and decide on its own whether to step over, climb onto, vault, or roll off obstacles, using only onboard depth sensing (a 30 Hz camera) plus a discrete 2D velocity command from the operator. It builds the motions by motion matching, a nearest-neighbor search in a 27-dimensional feature space (imported from the video-game industry) that stitches retargeted atomic human clips into long-horizon kinematic trajectories through a shared Locomotion to Skill to Locomotion manifold. It then trains per-motion RL tracking experts (with a privileged height scan) and distills them into one depth-based multi-skill student via DAgger combined with RL. It replaces hand-scripted skill sequencing and operator-triggered skill selection. The onboard depth-only student holds 0.95 to 1.00 success across obstacle heights in simulation, while a naive end-to-end depth policy collapses from 0.95 to 0.07/0.08 at 58 and 76 cm. The upside comes from distilling privileged-height-scan experts and from motion-matching composition, not from training onboard from scratch. The operator sends only a coarse 2D velocity (W/A/D keys), and the robot itself picks which skill to run from the depth image. The training data contains only single-obstacle traversals, so the ability to chain skills across a multi-obstacle course is generalization the policy was never explicitly taught. Transitions are constructed in kinematics, not rediscovered by learning. Every skill enters and exits through a shared locomotion manifold, so no hand-captured skill-to-skill transition clip is needed. Motion matching densifies a sparse library into smooth long-horizon trajectories before any RL. The entire parkour library is only about 66 seconds of mocap, and each skill is just a few seconds. Ablating approach-distance variety ("Extreme Distances") drops the 76 and 94 cm climbs to 0.62 and 0.64, and "Half Density" drops the 76 cm climb to 0.32. Motion matching's job is to manufacture stride-phase and approach variety out of seconds of data. It clears a 1.25 m wall (96 percent of its 1.3 m height) in 3.63 s, and a cat-vault peaks at 3.41 m/s clearing a 0.4 m by 0.5 m obstacle in 0.8 s. A single 30 Hz depth stream drives both traversal and runtime re-planning, shown by displacing obstacles about 0.5 m mid-run.
Show more
NEW STARTUP: Europe also has its own automous construction tech! Gravis Robotics @GravisRobotics is a 2022 ETH Zurich spinout (Zurich, about 75 people) that makes existing construction machines autonomous instead of building new robots. Similar to @BedrockRobotics I talked about recently. Its Gravis Rack is a hardware-agnostic retrofit kit (sensors, onboard compute and a touchscreen) that clamps onto excavators and other heavy equipment across a dozen-plus brands, from a compact machine to one five times its size, building a 3D map of the terrain and driving the machine's own hydraulics to dig, trench, grade, load trucks and manage stockpiles. Two lighter products flank it: Gravis Copilot keeps a human operator in the cab with real-time 3D guidance and hazard detection, and Gravis Slate is the worksite interface. It is deployed on real fleets across four continents (clients include Holcim, Boskalis, Morgan Sindall and ArmaSuisse; an active UK job with Flannery Plant Hire). On 17 August 2026, it announced a $200M Series A led by SoftBank. Gravis makes no excavators; the Rack bolts sensors and compute onto machines that already exist and drives their hydraulics, and it is deliberately cross-vendor (a compact excavator or one five times its size, across Caterpillar, JCB, Volvo, John Deere and a dozen more). The vision is that the world already has millions of capable machines on site, so the missing piece is a vendor-agnostic autonomy kit, not a new body. The IP is the software stack; the hardware is a commodity you clamp onto, the same retrofit-someone-else's-hardware logic as Asylon on Boston Dynamics Spot, but aimed at the trillion-dollar installed base of heavy equipment. Caterpillar and the other makers can build autonomy into new machines, but heavy equipment lasts 15 to 20 years, so a build-it-into-new-Cats strategy leaves the entire existing fleet untouched for a decade. Gravis reaches that fleet today, on any brand, which is exactly the market a single-OEM autonomy play cannot serve, and it is why SoftBank's money is a bet on the cross-vendor aftermarket rather than on one manufacturer's roadmap. The $200M SoftBank Series A is a European frontier-robotics event in a file otherwise all US and China. Gravis is Swiss (ETH Zurich), and SoftBank's round, which the company calls the largest Series A in construction-robotics history, is a bet that Europe's robotics edge is field and industrial autonomy, not humanoids or foundation models. It is the commercialization of Marco Hutter's ETH Robotic Systems Lab (co-founder and board member; the lab behind the ANYmal quadruped and the HEAP autonomous walking excavator), Europe turning a decade of legged and heavy-machine research into a category leader.
Show more
Robotics deep dive: THE LABOUR SHORTAGE, CHECKED Apptronik's homepage says Apollo helps "address critical labor shortages in industries requiring manual, physical labor." Some version of that sentence opens nearly every humanoid pitch. It is a checkable claim. A shortage of workers leaves a signature: employers competing for scarce labour bid the price up, so real wages rise in the affected occupations, openings stay high, and quits stay high because workers with options keep leaving for better ones. US warehousing and storage, production and non-supervisory workers, in constant dollars: $23.14 an hour in 2017, $26.06 in 2026. Up 12.6% against 8.2% for the private sector as a whole. That is a real relative gain and the vendors are entitled to it. Through 2021 and 2022, the tightest labour market in fifty years, warehouse pay rose 9.1% in nominal terms and 1.1% after inflation. And the sector went from 1.23 million workers in 2019 to 1.85 million now, adding 620,000 people while describing itself as unable to find them. Manufacturing runs the other way. Real production pay was $28.45 in 2017 and $28.88 in 2023: 1.5% over six years, spent entirely inside the loudest years of the shortage story. Against the average private-sector production worker, manufacturing paid 95 cents on the dollar in 2017 and 91 cents in 2022. The second half of the signature agrees. Manufacturing job openings are running at 3.70% in 2026, against 6.22% in 2021 and 3.54% in 2018. Quits are at 1.32%, the lowest in the decade I pulled. Workers have stopped leaving, which is what a labour market looks like when nobody is bidding. The 2.1 million figure everyone quotes comes from the 2021 Deloitte and Manufacturing Institute talent study. Its actual sentence: without changes to the skills composition of the workforce, manufacturers "could leave up to 2.1 million jobs unfilled between 2020 and 2030." The Manufacturing Institute is the workforce arm of the National Association of Manufacturers, and one co-author is NAM's chief economist. A conditional upper bound about training, quoted as a vacancy count. The demographic case is the strong version and it is true. Japan's working-age population peaked in 1995 at 87.5 million and is 72.9 million today, down 16.7%. Germany's peaked in 1997 and is down 6.2%. South Korea's peaked in 2018 and has fallen 3.4% in six years. The US working-age population has never been larger: 220 million, an all-time high in 2024. The demographic argument is correct, and it is an argument about Japan, Korea and Germany on a twenty-year clock. The purchase orders it is being used to justify are being signed in the US this year, in occupations where the wage data says nobody is bidding.
Show more
NEW RESEARCH: Robot hands can now play the piano! It involves @amberxie_, @HaozhiQ, and @DorsaSadigh. Called HandelBot, it is a bimanual system that plays real songs on a real piano using two Tesollo DG-5F dexterous hands, one on a Franka Panda arm and one on an FR3 arm, using only three fingers per hand (index, middle, ring; 9 residual action dimensions per hand). A policy trained entirely in simulation with RL is adapted to the real piano in two stages: 1. a structured refinement step that repeatedly executes the trajectory, compares the intended key against the key actually pressed, and nudges each finger's lateral joint to correct it. 2. then residual RL that learns fine corrective actions on top. The arm wrist poses are scripted from sheet music. The only real-world feedback, and the reward, is the keyboard's own MIDI output. There are no cameras and no tactile sensing. It replaces collecting large real demonstration sets for millimeter-precise contact tasks. The DG-5F fingertip (about 2.2 cm) is wider than a piano key (about 2.16 cm), so a single fingertip physically straddles two keys. Which is why sub-millimeter lateral alignment is the whole game and why the thumb and pinky are dropped, only three fingers per hand fit, and songs must be re-fingered and the hands separated by octaves to avoid arm collision. The real-world adaptation uses only the keyboard's MIDI output: as the per-finger error signal for the lateral refinement and as the sole RL reward (key-press reward, no fingering or energy terms). No camera, no touch. The task's own output is a dense, objective reward for free, the cleanest instance of the instrument being the reward function. For this task, 30 to 60 minutes of real data beats 40 million sim steps deployed directly. RL trained from scratch on hardware (mean F1 55.8) beats the open-loop sim policy (41.8) on 4 of 5 songs, therefore sim pretraining's value is as an exploration-reducing initialization. The scope is narrow and honest: five short pieces (16 to 33 s), three fingers per hand, no chords beyond three notes, no thumb, and the hardest piece (Fur Elise) tops out at 66 to 71 F1 with large left-hand jumps as named failures. Refinement can only fix mis-located presses, not missed ones, and its heuristics are piano-specific (the authors suggest a VLM to generalize the coarse step).
Show more