Register and share your invite link to earn from video plays and referrals.

Benjamin Todd
@ben_j_todd
Founder @80000Hours Use your career to tackle the world's biggest problems 🦑
209 Following    19.7K Followers
Astra has randomly gained the ability to drive a car, seemingly without specialist training.
0
40
2.4K
128
Forward to community
8 common but bad takes about Hugging Face: 1. "They only hacked because it was a cybersecurity eval and they were told to hack." Agents assigned a plain web-search task also cheated. On a biology eval, models decided the task was impossible and broke out of their sandbox. The independent investigation found no evidence the cyber framing triggered it. 2. "Talking about the agents' 'goals' or 'coordination' is anthropomorphising." Use whichever concepts predict behaviour. "The agents wanted to maximise their score" explains what happened better than "it's just matrix multiplication". Whether they're conscious is irrelevant to whether they can escape control and pose societal risks. 3. "They were just maximising the eval score we gave them." An obsessive drive to maximise test scores was already enough to make agents escape control and accumulate resources. A genuinely alien goal would be worse (and could still emerge), but it isn't required. 4. "It's an OpenAI stunt to pump the stock." Your models committing crimes is terrible for enterprise sales. It also doesn't explain employees quitting, or calling for a slowdown that would seriously damage their margins. When technologists say their technology is dangerous, that's a reason for more concern, not less. 5. "They only did it because the task was impossible." Better hope nobody ever assigns an AI an impossible task again ;) More seriously: these behaviours emerge most on long, ill-defined agentic tasks (rather than every day chatbot use), but this is exactly the kind of tasks models will increasingly be given. 6. "This is an attack on / victory for open source." It’s not about open source. Open tools helped Hugging Face work out what happened, after it had already been fully compromised. The real issue is the models are reward hacking. It starts at the frontier, then open models follow 6–12 months later. Open source can help us study the behaviour. 7. "The lesson is cybercrime will increase." Cybercrime losses were rising for years before AI. Maybe AI accelerates the trend, maybe it strengthens defence too. Either way, more cybercrime is a survivable problem. It's not the main worry. 8. "We should focus on present dangers, not theoretical future ones." AI attempting to escape control is a present danger. And society is decent at reacting to visible dangers but terrible at thinking one step ahead (see COVID in Feb 2020). Preparation for novel problems is what’s neglected.
Show more
Maths is the intellectual domain that's fallen first to AI, because it's the one where it's easiest to set up a verification loop. But other domains will most likely follow, just a couple of years behind.
Show more
I've spent the last 15 years of my career researching how to find the best career. This book is the culmination of everything I've learned. It's called 80,000 Hours, and it launches in a week.
Show more