Register and share your invite link to earn from video plays and referrals.

Search results for VLM
VLM community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including VLM
Transferring the Intelligence of VLMs to Robotic Control paper:
We comprehensively evaluated 16 recent frontier VLMs - including Opus 5.5 and GPT-6 Sol/Luna - on whether higher effort led to better document parsing performance. Higher effort typically leads to improvements on other benchmarks (coding, knowledge work), but up until recently it wasn't obvious that this natively improved capabilities for reading PDFs. Results: ✅ Out of the frontier models, Opus 5.5 has the best performance relative to its price. It's especially good at parsing tables. ✅ Astra is also quite good, but starts at a more expensive price than Opus. ✅ GPT-6 Luna is more compelling at the cheaper end of doc parsing If you're parsing documents at scale, you'll still want a dedicated OCR solution like LlamaParse ( that has better performance at a cheaper price. But if you're parsing docs "in the agent loop" within an app like Codex/Claude code, and you're too lazy to integrate a dedicated solution, then Opus 5.5 is currently the leader. Full results on ParseBench:
Show more
“In-Context Robot Learning with VLM Agents” Most robots need lots of task-specific training before they can do something new. This paper shows you can instead just show a general VLM what to do and let it adapt directly from context. GPT-Policy essentially turns robot learning into prompting. So they gave the frozen VLM an example of the task, even just a human video with no robot action labels, and it can translate what it sees into robot actions on the fly. This brings few-shot learning into the physical world, where teaching a robot a new behavior could look more like showing it an example than retraining a policy.
Show more
🤖 What if you could pilot a robot with a VLM that never sees a single robot training example? RoboDawn tackles exactly that question. Title: Transferring the Intelligence of VLMs to Robotic Control (RoboDawn) URL: It gives a frozen, pretrained VLM a human-intuitive interface of translation, rotation, and gripper commands, plus a handful of in-context demonstrations, and lets it directly drive a robot arm. Three things stand out. 🎮 A game-like control interface The VLM issues discrete move, rotate, and gripper commands, all defined relative to the gripper interaction point. This lets it reuse spatial manipulation knowledge it already picked up from web-scale pretraining. 📚 One demo makes a huge difference No parameter updates at all. Just a command primer plus task demonstrations as context lift success on RoboTwin 2.0 from 53.2% zero-shot to 73.6% one-shot. 🏆 It beats robot-trained policies outright With zero task-specific training, RoboDawn's zero-shot performance already surpasses policies trained on dedicated data, like π0.5 (46.0%) and LingBot-VLA (50.4%). On a real Franka robot it hits a 90% success rate. It suggests the real bottleneck may not be collecting more robot data, but designing the interface that unlocks the intelligence VLMs already have. #Robotics# #VLM#
Show more
D Mitja Ilenic returns to NYCFC from loan
Sensadimes are good for the soul 🧘‍♂️