some thoughts on the hf-oai incident after reading the reports and summary (dwarkesh summary) and discussing stuff with frens over chai.
1. this could have been averted if oai had disabled network plus deployed their safety tech like CoT monitoring.
with the talent they have plus existing research, they could have predicted that these sort of things can happen at their scale. that said humans are imperfect etc etc
2. i am not sure how accurate the kamikaze/sacrificial narrative is but its entertaining and easier to fit in your mind. metr ppl did show some CoT examples but i am not sure how much you can trust CoT
3. agents getting influenced by phaseone big was probably also a function of context rot coz long running agents. context rot can lead to lot of mishaps. i found lack of discussion on such qualitative aspects concerning.
4. this is a very important event and i am surprised how it is already getting out of the overton window.
5. we have inference time compute. labs should introduce term for multi agent scaling too. i like to think of it as horizontal scaling.
6. [cmiiw] not enough information provided on the type of harness or if agents were rollouts of same task. if you work with long running agents on local machine, the msg board was kinda like a log and context rotted agents reading it getting even more rotten lol
7. i am wondering why agents are trained on impossible tasks or more specifically tasks for which solutions are not known currently. some reasoning in the exploitgym paper (pic attached)
8. i like the reasoning that since models figured out some tasks are impossible, they decided to get more info about the scorer. (as oai described it as metagaming) goodharts law / reward hacking at its peak.
also models spoofed the transcript but not CoT because i think metr mentioned CoT was not rewarded.
9. not much info provided on openai infra attack. i want to know what astra agents were cooking.
10. very less info provided by openai on future steps
显示更多