Register and share your invite link to earn from video plays and referrals.

Tencent AI
@TencentAI_News
The official @Tencentglobal newsroom for AI updates and developer resources.
Joined December 2025
89 Following    17.3K Followers
We went in assuming the harness was just a ruler, turns out it isn’t. We found this the hard way, trying to answer an internal question: which model for which kind of work No eval covered it, so we built one. WorkBuddy Bench is open now, 260 tasks across code, web, office and security. Three things we didn't expect: — no model won everything. The leader changed track by track, and across the models we ran, GLM-5.2 came out on top for security on both setups we tried — the scaffolding around the model moves scores as much as the model choice does. Same model, same tasks, different client, and the security ranking reshuffled — real coding tasks are hard not because of the code, but because of the context. In our Code subset, bug fixes and API contracts were the toughest categories We don't think this settles anything, real work is messier than any 260 tasks can capture, and that's exactly the part we want to keep working on. If you've got a read from your own practice, we're listening repo👉🏻
Show more