가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Léo
@LeoKharon
Robotics research & updates. Co-host @roboticsstack, the weekly pod on what's actually shipping in 🤖
가입 November 2022
469 팔로잉 중    1.9K 팬
Robotics deep dive: GEMINI ROBOTICS 2 Google DeepMind shipped its new robotics stack, here is a rundown: it is not one model but three, split by job. 1. Gemini Robotics 2 is the body. A vision-language-action model for direct control: an image and an instruction go in, robot actions come out. The headline is whole-body. For the first time one policy drives a full humanoid, feet to fingertips, plus dexterous hands and grippers. Private preview. 2. Gemini Robotics ER 2 is the brain. A vision-language model for embodied reasoning, built on Gemini 3.5 Flash. It takes text, image, video and audio, up to a 128k context, and outputs text. Its job is to plan multi-step tasks that run for minutes, orchestrate tools, and coordinate several robots as a team. It thinks about the next step while the body is still moving. In preview in the Gemini API and AI Studio. 3. Gemini Robotics On-Device 2 is the edge. A lightweight VLA that runs locally on the robot, for sites with no reliable cloud. Inputs are text, images and proprioception, outputs are action values. The striking part is adaptation. It fits a completely new robot, a different shape and different sensors, in a few hours on typically fewer than 200 examples. For trusted testers. This is the division of labor the field keeps converging on: a slow reasoner on top, a fast action model below, and a local fallback on the hardware itself. Google now ships all three under one name.
더 보기