Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
need to correct this. I spoke to Transluce after posting and they said it's not clear in the evidence they have that the Sept. 16 activity is OpenAI. It could be a different source.
There's 2 major new things we learned from the Australia govt disclosure and Transluce report
- It's not just agents tasked with cyber tasks that end up hacking
- Lastest activity was Sept. 16, showing OAI has struggled to lock down all the agent activity
I think a good job to tell your kids to go into is something like a genetic counselor or patient advocate. Where you're doing the hard work of explaining medical processes and decisions but you are not doing the diagnosis yourself
@cwarzel the radiologist argument is starting to break down. Of course at some theoretical point we stop needing people with 10 years of schooling to do patient relations if the machine is doing 100% of the diagnosis
@cwarzel the radiologist argument is starting to break down. Of course at some theoretical point we stop needing people with 10 years of schooling to do patient relations if the machine is doing 100% of the diagnosis
the core disagreement in the pod is essentially that Jensen is not AGI-pilled and Ezra is. Jensen thinks it will be a world-changing technology but insists that humans will always understand every level of it, which to people in the AI labs is a total contradiction
Some NPR podcasts started getting mysterious comments on Spotify. They made no sense to the staff reading them – until someone from a younger generation cracked the code.
Hear the story:
Five cops in Indianapolis were just criminally charged with abusing Flock to spy on their ex-wives, their wives' ex-partners and women they met on duty. Audit was triggered by us at the @washingtonpost, the prosecutor said. Story incoming.
Scoop: The White House has asked OpenAI and Anthropic not to share their AI models with the U.K. government’s testing agency until the models have gone through U.S. testing.
“Because they’re American companies and this has been our policy with every new frontier model that comes out,” the senior admin official told me.
WH wants this sequence: U.S. review -> secure U.S. systems -> share models with U.S. partners
Anthropic appears to have complied but OpenAI has not said if it will.
w/ @JoeBambridge1
There's 2 major new things we learned from the Australia govt disclosure and Transluce report
- It's not just agents tasked with cyber tasks that end up hacking
- Lastest activity was Sept. 16, showing OAI has struggled to lock down all the agent activity