The Braintrust eval library has a repo of skills your coding agent can read, so you can easily build and run evals on your data.
Here's an example of using a skill in Claude Code to compare the Codex CLI and Pi, both running GPT-5.6 Sol, on a 30-task stratified SWE-bench Verified dataset. The skill saves those experiments and traces to Braintrust.
Try it yourself →
See more skills →
顯示更多