Jev from
@typesafeai can replace your LLM-as-a-judge for scoring agent responses.
It returns a choice or numeric result with information about uncertainty, so you don't need to spend time and resources prompting a general-purpose model into an LLM judge.
Use Jev as a judge scorer in Braintrust and review its selected answer, confidence, and probabilities alongside the score. Trace Jev calls from your own application with the JavaScript or Python SDK.
Read more →