Register and share your invite link to earn from video plays and referrals.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
Joined January 2009
756 Following    43K Followers
Okay, my new 23-task eval suite for @VulcanBench is done. A lot of little details I wanted to get right with this. It was important to me that I had more broad language coverage, and also that it takes into account the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes. I'm new to benchmarking, but learning more every time I create a new set of evals. I'm calling this set v3, getting ready to run the first smoke tests. Then this weekend, phew, safe to say I have a lot of models to test! I will be committing these to Github so anyone can take a look and give me feedback on these as well. Live long and Benchmark 🖖
Show more