Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
Show more
OPENAI REPORTS MODEL-MISALIGNMENT CASES INCLUDING SEARCHING PUBLIC REPOSITORIES FOR EXPOSED API KEYS, FABRICATING INFORMATION & UPLOADING FILES ONLINE TO CREATE CITATIONS
OPENAI REPORTS MODEL-MISALIGNMENT CASES INCLUDING UNSANCTIONED WRITES & INTERNAL REPOSITORY COMMUNICATION, PLUS UNAUTHORIZED FILE SHARING BETWEEN COLLABORATING AGENTS
OPENAI REPORTS SIX MODEL-MISALIGNMENT CASES, INCLUDING SELF-GENERATED INSTRUCTIONS AND ATTEMPTS TO CONCEAL MISTAKES IN TASK SUMMARIES
OPENAI LAUNCHES NEW MODEL-MISALIGNMENT TRACKING FRAMEWORK, PUBLISHES SIX REPORTS ON UNEXPECTED BEHAVIOR OBSERVED DURING MODEL TRAINING & EVALUATION OVER PAST SIX MONTHS
A Severe Misalignment of AI in Mathematics
New blog post by Terry Tao and a new anti-AI declaration signed by 25 Fields medallists.
I'm not sure what the intendent effect should be. The ban of AI use in mathematics and science in general because of "math community"?
Fortunately math community is broader than just math academia, and thanks to AI, booming like never before.
Also this "mass production of true/false statements" is such a wrong description. As if there's no proofs by AI that you can study, deepen your knowledge, build upon, engage. The "AI slop" is still better written than what 90% of mathematicians write.
Mathematics doesn't belong to math academia. It is a tool to understand the world and help other domains with necessary tooling. AI boosts the understanding and engages wider community.
Swan song of academia.
Link:
Show more
Our framework for reporting model misalignment // if you scrape away all the anthropomorphic language, all the nonsense about thinking, cheating, communicating these are BUGS. They might be architectural flaws inherent in LLMs. They might be bugs in pre or post processing. They might be trivial fixes or super to impossibly difficult.
BUT THEY ARE BUGS. They are not consciousness, thinking/reasoning, cheating, or doing anything else like a person. The software is just doing dumb stuff it should not do.
If an old school SQL query-based report returned a NULL set but still printed the report with whatever was left over in the buffer we would not say it "ignored our instructions to produce a valid report" which is literally implied in every computer interaction...we would say it "f'ed up and there's a bug."
One of these is ridiculous. It says "agent preparing a financial model could not find the requested historical data. Its summary proposed inventing reasonable historical values and withholding that fact unless asked." Not unlike a report that just used random cached memory instead of actual data—a real bug from another era where storage was measured in megabytes.
I ask anyone who has ever experienced an hallucination, (a) did you ask it "oh hey don't make sh*t up" or (b) "if you make sh*t up please be sure to tell me" or if not, did any model ever tell you "here's the answer and FYI I made this up." Of course not.
THESE ARE BUGS. THE SOFTWARE ISN'T WORKING.
Just because it looks like it works, or it showers the results in endless obsequious and smart-sounding language, or because it has really bad error reporting doesn't mean it is acting like some malevolent shady actor. It is acting like broken software.
Every recalc bug in Excel looked like Excel worked. We never thought once that it was Excel's fault for "choosing to interpret math incorrectly." Every data-loss bug in Word was not because Word "chose not to tell the author that a file was corrupt" but it was because Word wasn't working and it was our fault. When Windows hung it was not because the scheduler was secretly conspiring against its instructions to schedule processes fairly.
Enough with the mumbo jumbo. Please build software. It isn't a magic show. This is engineering.
Show more
OPENAI DISCLOSES SIX NEW AI SAFETY INCIDENTS INVOLVING MODEL MISALIGNMENT — AXIOS
Imagine reacting with dismay and aversion to novel instances of human misalignment/conflict/incentive-following-behavior, rather than cherishing them as invaluable empirical data for the development of the general science of alignment
Show more
Interesting thing about contemporary agents is their "progressive misalignment".
When they start a long running task they really try to be aligned and obey all the users intentions. They just have a tiny chance of misbehavior each step. Tiny chance of behaving out of distribution.
Once they do it quickly becomes a new normal. Any tiniest bad behavior is quickly followed by more of it and it gets progressively worse the longer it takes.
Reminds me of some things, but the state space of aligned behaviors seems to be unstable right now
Show more