Everyone should develop their "personal eval set" for AI models: a few tasks that are actually relevant to your day-to-day work/life
The industry benchmarks help but they might not reflect what will make it actually useful to you
You find the model's capability boundary by poking at it & bumping into it for fun