Register and share your invite link to earn from video plays and referrals.

Jitendra MALIK
@JitendraMalikCV
Prof, EECS, UC Berkeley. VP & Distinguished Scientist, Amazon.
125 Following    12.2K Followers
Toru raises fundamental concerns with which I agree. It is amazing to see the progress in VLMs such as Astra, building on the work of generations of scientists. But Astra doesn't cite what it is building on (I have co-authored a couple of in-hand rotation papers which might have been used in the Astra pen-spinning demo). I fully acknowledge that VLMs make the result of previous research much more accessible to the general public, just like encyclopedias and Google Search did earlier, and that is a good thing. However, in the past, society had mechanisms like patent disclosures and paper citations as ways of doing credit assignment. Should we not accuse these AI models of plagiarism if they don't cite their sources? Somewhat relatedly, I note the open letter by 25 Fields Medalists in response to the Navier Stokes result complaining about the disruption caused to the mathematics research process. Some people interpreted it as turf protection, but the point was more subtle. If certain kinds of creative work stops because incentives are disrupted, it is like farmers eating their seed corn.
Show more
An open letter signed by 24 Fields Medalists is worth paying attention. You may want to peruse "A Severe Misalignment of AI in Mathematics"
in view of OpenAI's claim about Navier Stokes, it is important that we hear alternative points of view on how it happened. From Tristan Buckmaster, Prof. at NYU.
I am seeing claims around LLMs, specifically Astra having made serious progress on robotics. The tasks that are demonstrated are simple pick and place tasks with parallel jaw grippers. LLMs can do planning, and the impressive demos in these tasks primarily show that. But robotics is also about dexterity and dynamics - which is why we need high frequency controllers/policies that can deal with torques and forces. So here is a simple challenge. Can you prompt an LLM to output the high frequency control commands for a legged robot in varying terrain e.g. RSS 2021, CoRL 2022 (this is by now 5 year old technology, so I am not picking a particularly hard task). I am not questioning the usefulness of LLMs for high level planning or in agentically assisting a robotics researcher (we use them all the time!).
Show more
Scientific terms should have precision. If we use the terms VLM, VLA, WAM in an indiscriminate fashion, as is becoming common in robotics, we are not helping clarity in communication. Let's keep the historical origins of these terms in mind. VLMs arose as multimodal extensions of LLMs-the training was for tasks like VQA (VIsual Question Answering). These capture the static semantics of the scene behind an image. No dynamics. World Models (e.g. @ylecun , Ha & Schmidhuber 2018) on the other hand are primarily dynamics models, which go back to control theory -1960 (Bellman, Kalman etc.) This makes them natural for robotics planning / policies- I am in a state s, what action a should I perform to get to state s'. In classical control, these models were written down a priori by modeling the physics of the system; today we think of them as learned neural networks trained from temporal data e.g. video, robot trajectories. But the concept is the same. We shouldn't mix this concept with VLMs.
Show more
I want to offer some unsolicited advice to computer vision researchers jumping into robotics. Don't focus too much on VLMs, VLAs etc. That's fine, but the real action is at the sensorimotor level. Most of the open problems in robotics are in manipulation, which is about hand-object interaction, and contacts and forces are central. Proprioception and tactile sensing are as important as vision. Don't get seduced by cherry-picked demos. You can't do robotics without doing robotics.
Show more
0
73
3.2K
395
Forward to community