註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Braintrust
@braintrust
Active observability for agents in production.
加入 August 2023
60 正在關注    7.7K 粉絲
We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more →
顯示更多