Register and share your invite link to earn from video plays and referrals.

Kai
@kaicathyc
Occasionally here. Currently research @OpenAI.
111 Following    23.9K Followers
> The best model is no longer the unethical one.
Improvements for Astra came from more general techniques in development long long before the Hugging Face incident. ExploitGym Honeypot was very recently added as an eval following that incident and is out of distribution for our RL runs. There are also clear improvements across a wider range of behaviors in deployment simulations, deception evals, and realistic computer-use tasks, though these still reveal failures and substantial room for improvement. We should have definitely done a better job explaining where we think the alignment improvements came from in the system card and that was a miss. Measuring alignment generalization is a core part of our research program - we do not benchmark-maxx alignment evals. This would be horrendously stupid and I hope no lab is doing this. That being said, metagaming and eval awareness are real challenges for any effort to measure alignment, including ours, and understanding their effects is an active focus. We're spending a lot of effort to improve our eval techniques and study worst-case alignment and metagaming behaviors in our models.
Show more
Yes the blog is sexy but also highly recommend reading through the system card. We share a lot of evals and analysis we ran on the alignment of this model and discuss continued challenges for both alignment and monitoring.
Show more