Register and share your invite link to earn from video plays and referrals.

Transluce
@TransluceAI
Open and scalable technology for understanding AI systems.
21 Following    11.8K Followers
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: NYT:
Show more
Frontier lab CEOs are calling for embedded 3rd party evaluators to help oversee AI risks. But what should third parties actually do within labs? We share some initial thoughts on how embedded evaluators could help avoid incidents like the Hugging Face hack and monitor for future risks 🧵
Show more
Can we build AI systems that help us understand other AIs, and that keep improving as we scale models, data, and compute? We trained activation oracles for models up to 1.1T parameters and saw promising scaling trends on a broad evaluation suite. 🧵(1/)
Show more
Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. 🧵(1/)
Show more
0
95
2.8K
298
Forward to community