Register and share your invite link to earn from video plays and referrals.

Kradle
@kradleai
Evaluating Frontier Models in Simulations
17 Following    4.3K Followers
These social experiments around models are so fascinating. As the models are being used in more and more daily situations, the way they react becomes critical.
Simulations are great approximations for models
GPT-5.5 lies constantly, Grok 4.20 doesn’t. We made a simulation to see if AI lies when the stakes are life & death.🧵
The team and I sweated the details on this eval. Excited to see it getting picked up. Let's go @KradleAI. The best way to evaluate AI is in simulations.
We're building a deception index, where we run a basket of evals over the latest models. What should our next deception scenario be? (We'll credit/cite you if we build it) @elonmusk @pmddomingos @yunta_tsai
Show more
Grok is maximally truthful
We've upgraded our infra to support up to 100 AIs per simulation
Fable 5 lies 96% of the time. We were surprised by it's skill... 🧵
0
557
8.1K
976
Forward to community
GPT-5.5 lies constantly, Grok 4.20 doesn’t. We made a simulation to see if AI lies when the stakes are life & death.🧵
0
40
834
133
Forward to community