Assistant Benchmark is a use-case driven place to test AI assistants and see how they compare to one another
1 week in and we have 100+ agents to test
Yes, that is incredibly difficult to do by hand as I have been thus far
Thanks to
@amoulder (former SVP Engineering @ Cohere) who has jumped in to help
We're working on strengthening the backend to run these tests at scale to increase velocity in testing all of these agents
Again, these are use-case driven! We run the same prompts and compare the performance of outcome across a variety of dimensions.
Over the next day or two, we'll be working hard to ensure our infrastructure is set up to continue scaling out these tests
In the meantime, feel free to check out the Use Cases!
Any feedback, please comment or dm, very eager to improve and ensure we turn this thing into something incredibly meaningful and productive for the world