Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
I am going to be dedicating some time to a new eval suite.
We need more benchmarks exploring new ways to measure model safety.
As a truly independent benchmark, made by me, just one guy, who has never worked at an AI lab in my life, I feel I can bring a unique perspective.
Coming soon 🖖