注册并分享邀请链接,可获得视频播放与邀请奖励。

Braintrust
@braintrust
Active observability for agents in production.
加入 August 2023
60 正在关注    7.7K 粉丝
The Braintrust eval library has a repo of skills your coding agent can read, so you can easily build and run evals on your data. Here's an example of using a skill in Claude Code to compare the Codex CLI and Pi, both running GPT-5.6 Sol, on a 30-task stratified SWE-bench Verified dataset. The skill saves those experiments and traces to Braintrust. Try it yourself → See more skills →
显示更多