注册并分享邀请链接,可获得视频播放与邀请奖励。

Xiaokang Chen
@PKUCXK
Researcher in Multimodal Team @deepseek_ai Previously Bachelor & Ph.D at Peking University @PKU1898 Opinions are my own.
加入 June 2022
63 正在关注    6.4K 粉丝
Absolutely mind-blowing work. 🤯 This reminds me of when we were exploring the concept of 'thinking with visual primitives'. We noticed the model excelled in synthetic scenarios, but real-world generalization for complex tasks still faced bottlenecks. Back then, I kept thinking: "If only someone could provide manual annotations for these real-world scenarios..." Using point-form visual primitives for reasoning would be powerful for tasks like state tracking, spatial topology, and logical testing, but they may heavily rely on high-fidelity, real-world labeled data. I think it’s the same across other domains too. The journey toward AGI cannot skip the hard, grounded work of human-labelled data. Massive respect to this effort! 🫡
显示更多
@skalskip92 You might not believe it, but I simply manually annotated over 2,000,000 human body parts with ultra-precise detail. Probably no one else could do that.