We went in assuming the harness was just a ruler, turns out it isnโt.
We found this the hard way, trying to answer an internal question: which model for which kind of work
No eval covered it, so we built one. WorkBuddy Bench is open now, 260 tasks across code, web, office and security.
Three things we didn't expect:
โ no model won everything. The leader changed track by track, and across the models we ran, GLM-5.2 came out on top for security on both setups we tried
โ the scaffolding around the model moves scores as much as the model choice does. Same model, same tasks, different client, and the security ranking reshuffled
โ real coding tasks are hard not because of the code, but because of the context. In our Code subset, bug fixes and API contracts were the toughest categories
We don't think this settles anything, real work is messier than any 260 tasks can capture, and that's exactly the part we want to keep working on. If you've got a read from your own practice, we're listening
repo๐๐ป