註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Xiaokang Chen
@PKUCXK
Researcher in Multimodal Team @deepseek_ai Previously Bachelor & Ph.D at Peking University @PKU1898 Opinions are my own.
加入 June 2022
63 正在關注    6.4K 粉絲
You can try the following two prompts in the Thinking mode (via web/app) to get a better model experience in certain domains like counting (Note: keep a line break after the bracketed titles):: [Think with Grounding] ....... [Think with Pointing] ...... These two prompts encourage the model to adopt bounding boxes or points (which are classic fundamentals in computer vision) in its thought process. Personally, I love the pointing approach for solving abstract topological/reasoning tasks. Using points to represent continuous trajectories makes the MLLM's reasoning process feel much more human-like. Speaking purely from my personal exploration: Getting a multimodal model to accurately represent continuous trajectories with points is still a highly challenging frontier task for the entire industry. The current performance on real-world scenarios still has a long way to go.
顯示更多
Vision is now live on web and app. 👀 Come test the new eyes, but give its pure text capabilities a try while you're at it.
0
9
180
20
轉發到社區