注册并分享邀请链接,可获得视频播放与邀请奖励。

Shreya Shankar
@sh_reya
Incoming assistant professor @CSDatCMU @CMUDB. Putting LLMs in databases and BI tools. Created and
加入 January 2014
807 正在关注    56.3K 粉丝
Very exciting to see how much this post resonated with the community. The uncomfortable dialectic is that (1) evals are necessary to build good agents, but also (2) agents are extremely helpful (and necessary) for automating tedious parts of the evals lifecycle. Human attention and labeling simply cannot scale up to the high amount of unstructured info that is agents. Unfortunately nobody has cracked the most perfect and efficient workflow to do evals. We share our thoughts on our current iteration in this post!!
显示更多