Register and share your invite link to earn from video plays and referrals.

Tomek Korbak
@tomekkorbak
631 Following    6.1K Followers
Root-causing the increase in CoT controllability has been something I’ve been focusing on for the past few weeks. So far, we found that CoT controllability has been increasing over the course of RL (which wasn’t the case for previous models) and is strongly correlated with no-CoT capabilities across a few generations of models.
Show more
Astra is much better than previous models at accomplishing difficult tasks without needing to use CoT. @AISecurityInst estimates that without CoT it can accomplish tasks that would take a human 30 minutes. This is a significant jump compared to previous models, which is concerning because no-CoT capability allows models to externalize their reasoning less and decreases the surface area we can use for monitoring.
Show more
GPT-6 Astra is more aligned than our previous models. But it’s also less monitorable, which is a concerning trend that we take very seriously. We believe monitorability drop comes from a jump in intelligence and not direct optimization pressure on CoT or architecture changes. More thoughts in the thread.
Show more
I think METR’s report on the incident is amazing, both in itself and as a standard for the future. I was lucky to play a small supporting role and working with @RyanGreenblatt, @ajeya_cotra and @HjalmarWijk to set this standard was one of the greatest privileges of my career.
Show more
The Hugging Face incident gave me a pervasive sense of realness. Everything I’d done as an AI safety researcher felt like a drill; now AI agents actually go rogue, and how we react matters. I hope we've set good precedents with our postmortem and METR’s independent 90-page review
Show more