AGENT EVALS is going to be a big thing and it’s a big white space to work on !!
I have been working with agents from the past 6-8 months now and I have been shipping it in production pipelines as well.
The thing about agent evals o have found out is -
They don’t exist !!
Yes, there is no proper eval that can be used in all the production grade agents or pipelines.
Almost every big company has a blog on this -
-Open Ai
-Anthropic
-Langchain
-Langfuse
-Data bricks
-IBM
-Hugging face articles
So many repos as well but all are scope bounded.
2 days back I had to score few checkpoints in my agentic harness.
How would I ?
Simply prepared 6-8 mathematical scores and combined them with some ground truth and that became my agent eval for judgement.
Now from the last 2 days any new model or pipeline or experimentation I do over those checkpoints, I can instantly scrape them on and off just via the scores.
Simple and easy.
Learn to build custom agent evals and you have a long way to go, the need for the same is gonna get even bigger !!
For resources check this research paper or research book as you may call it:
> GENERAL AGENT EVALUATION
顯示更多