王老师这篇我非常赞同。在任何模态上,评估都是被动的,反应式的(这是个bootstrape 问题)。无法指引或预测下一代能力升级"质变/涌现"的可能方向或障碍。甚至对当前模型能力的评估维度也是非常局限的。许多关键问题上是失效的。对模型在应用侧的实用性和有效性的评估也是不够的,训练和应用对不齐。
coding 和 agentic 跑得快也是因为评估好做,影响上下游的一切。
我最近两年的工作核心是做视频/图像/vwm 模型的评估,这个更难,更依赖人。测试框架也需要非常频繁地更新,追着参数量和需求端跑。
I’ve left Google DeepMind after an amazing chapter.
I’m incredibly grateful for the people I worked with, the things we built, and the lessons I learned from taking frontier AI research into production. DeepMind shaped how I think about research, product, evaluation, and what it takes to build AI systems at real scale.
As I wrap up this chapter, I wrote down something I’ve been thinking about a lot: evals.
We’re good at evaluating the models we have. We’re much worse at evaluating the models we’re about to build — especially if they cross into a new capability regime. We will have self-evolving models, but before that, we need self-evolving evaluations.
顯示更多