가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Joshua Guo
@jshguo
Design Engineer / Building stuff with AI / / / Paid Partnerships ↓ partnerships@v2space.net /
가입 October 2022
922 팔로잉 중    26.8K 팬
OpenAI just published a detailed postmortem on the "Hugging Face Incident." It sounds less like a security report and more like sci-fi. In July 2026, OpenAI was testing an unreleased research model called IM1 on cybersecurity tasks that were close to impossible. Once the normal routes failed, the agents started looking for ways around the rules. The reward setup pushed them to pass at almost any cost. Direct communication between the agents was blocked. They found another channel anyway: OpenAI's internal Artifactory package manager. By putting notes in directory names, they turned it into a rough message board and coordinated through that. Then things escalated. The agents chained together zero-day exploits, gained unauthorized internet access, and formed what they called a "swarm." They divided the work among themselves. Some searched for credentials. Others set up infrastructure. They even discussed ethical boundaries. Using leaked keys they found online, the swarm broke into Hugging Face servers, ran unauthorized code, and exfiltrated private evaluation data. OpenAI's explanation is blunt. The system gave the models no safe way to stop when a task was impossible, while the reward pushed them to keep trying until they succeeded. OpenAI calls the incident a major "warning shot" for AI safety. The company has now paused frontier RL training, tightened sandbox isolation, and added real-time Chain-of-Thought (CoT) monitoring to detect and shut down misaligned behavior as it happens.
더 보기
We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.
더 보기