登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Xiaokang Chen
@PKUCXK
Researcher in Multimodal Team @deepseek_ai Previously Bachelor & Ph.D at Peking University @PKU1898 Opinions are my own.
参加 June 2022
63 フォロー中    6.4K ファン
You can try the following two prompts in the Thinking mode (via web/app) to get a better model experience in certain domains like counting (Note: keep a line break after the bracketed titles):: [Think with Grounding] ....... [Think with Pointing] ...... These two prompts encourage the model to adopt bounding boxes or points (which are classic fundamentals in computer vision) in its thought process. Personally, I love the pointing approach for solving abstract topological/reasoning tasks. Using points to represent continuous trajectories makes the MLLM's reasoning process feel much more human-like. Speaking purely from my personal exploration: Getting a multimodal model to accurately represent continuous trajectories with points is still a highly challenging frontier task for the entire industry. The current performance on real-world scenarios still has a long way to go.
もっと見る
Vision is now live on web and app. 👀 Come test the new eyes, but give its pure text capabilities a try while you're at it.