Takeaways from the WSJ article about
@HacktronAI using Claude to get into OpenAI's monorepo and issue a pull request (before stopping and claiming their bug bounty)
- how many nation states have already broken in and gone much further and stolen a) algorithmic secrets and b) model weights or c) gotten access to user data; these are fair public interest questions
- how many have implants in / are dwelling in the openai network as I tweet this
- how pervasive is this level of softness to pentesting across all the labs and how far are the labs from the right operating point in the security and r&d friction trade space (probably pretty far it seems)
- given the hacktron folks used anthropic's models to pull this off what's the real public safety ROI of anthropic's cyber guardrails; they add friction for legitimate cyber defenders (like our developers at my startup!) but it appears with a bit of work you can use them in actual breaches as happened here
- as the article says, the wsj folks and
@S1r1u5_ had me do neutral technical review of the kill chain here pre publication; impressive from a human angle (
@HacktronAI reminds me of the best of my generation of hackers I looked up to as a kid!) but also from what the models can do; between this and the openai/hf thing, I'm emotionally in a place where I feel my world as a security person is being turned upside down in slow motion with respect to what's coming
- elite persistent hacking is becoming rapidly democratized. this is coming like a freight train. we need to harden the world's code and infra as fast as possible and today's non automated methods don't stand a chance of cutting it; the world is a soft target. I continue to be unsettled but very glad I left my comfortable job at Meta to do our automated posture hardening startup