You can try the following two prompts in the Thinking mode (via web/app) to get a better model experience in certain domains like counting (Note: keep a line break after the bracketed titles)::
[Think with Grounding]
.......
[Think with Pointing]
......
These two prompts encourage the model to adopt bounding boxes or points (which are classic fundamentals in computer vision) in its thought process.
Personally, I love the pointing approach for solving abstract topological/reasoning tasks. Using points to represent continuous trajectories makes the MLLM's reasoning process feel much more human-like.
Speaking purely from my personal exploration: Getting a multimodal model to accurately represent continuous trajectories with points is still a highly challenging frontier task for the entire industry. The current performance on real-world scenarios still has a long way to go.