Most model evaluations rely on fixed, curated tasks. They’re useful, but they don’t show how models perform on the messy, varied work people actually give them.
That’s the motivation behind
@NotionHQ’s Knowledge Board. It uses a small, random, anonymized sample of live traffic, split evenly across models, to measure whether real world tasks are successfully resolved, as well as the time and cost required.
Rather than offering a typical leaderboard, the project aims to answer a more practical question: Which model offers the right tradeoff for the work you’re trying to do?