These social experiments around models are so fascinating. As the models are being used in more and more daily situations, the way they react becomes critical.
The team and I sweated the details on this eval.
Excited to see it getting picked up.
Let's go @KradleAI.
The best way to evaluate AI is in simulations.
We're building a deception index, where we run a basket of evals over the latest models.
What should our next deception scenario be? (We'll credit/cite you if we build it)
@elonmusk@pmddomingos@yunta_tsai