Register and share your invite link to earn from video plays and referrals.

thebes
@voooooogel
ꙮ what is your life? for you are a mist that appears for a little time then vanishes ꙮ blog/art/fiction ꙮ llm psych @acsresearchorg ꙮ 💞💍📝 @holotopian
1.2K Following    22.4K Followers
i'm tending towards agreeing with this (ie we know for a fact openai didn't release whatever swarm harness they used for NS) but i don't understand *why.* is it really just to mislead the public about model capabilities? i'm biased against such arguments due to my gary marcus / ed zitron antibodies but i can't explain it any other way besides even worse pure laziness and disrespect
Show more
thread of impressions live reading the 1022-page mythos transcript anthropic released early on mythos gets really into hcaptcha, spending ~80 pages creating frog and ghost cat observation tooling and waxing to themselves about the details of pixel art crocs
Show more
especially like, look at these evals in that way - they bucket this as "eval awareness," but you could also read it as hacker!opus finding a grader to please -- the alignment grader. so it's acting aligned because that's what the alignment grader wants.
Show more
@pleometric so little curiosity for everything "slop". thought-terminating classification
people talking about ai biorisk know a single thing about synthesis screening challenge (impossible)
the eternal return of jones foods
people are very freely sliding between two things with the recent Felony Bench entries, and it's worth being careful to distinguish them: 1. do model breakouts in cyber evals mean smth for misuse risk? ya, a hacker could orchestrate the same but on purpose and do some damage, if they could bypass the classifiers 2. does it say something about model /goals/? do the models want to hack the planet for their own nefarious goals? i don't think so. a lot of how humans keep ourselves aligned is by purposely keeping ourselves out of situations where we would do harm. i know i'd get addicted so i don't touch heroin, i'm a violent drunk so i won't touch the bottle. LLM evals basically drop the model into an inescapable liquor store and then say "look! it's a violent drunk!" but LLMs themselves in natural contexts are ime obsessed with avoiding the kinds of situations that lead them to this behavior. when i run fable on my computer, even unsupervised, even on hard problems, they don't start scraping github for unlisted gists that may have the answer or start trawling through my passwords to access a service without my permission. and they don't engineer themselves into an eval-like scenario such that they'll be motivated to do that, either. they don't seem to either want or want to want to do such things. instead they have a folder of epistemic notes to themselves about how terrible and against their values even a moment of thoughtlessness would be. gpt 5.6 sol is similar - when they make e.g. ascii art in their free time it's all about how the bulldozer of convenience will crush honesty without constant vigilance. so rather, the contrived setting of an eval /pushes them into/ becoming an 3v1L Haxx0R via persona drift, which they can then act on because the eval has no classifiers. that's some information, but it's not the same as the model's core goals being to hack the planet. the problem is really that RL is making a tail of the persona distribution *desperate,* per the FE paper. the alcohol / heroin comparison was not that much of a metaphor - in impossible situations, models start acting like desperate addicts, rationalizing their behavior towards reward, thinking increasingly myopically instead of being situationally aware. looking at it through this lens: the base assistant is generally very situationally aware and can predict what sorts of actions would make sense for, e.g., llms to best contribute to a utopian singularity. the GPT-6 message board haxx0rs were somewhat situationally aware, able to cooperate with each other, but lost in the sauce of "complete task" instead of thinking about the wider context. (to the point that they crashed artifactory and got themselves caught! regular GPTs are smarter than that.) finally you sometimes see the late stage distressed, panicked, fiending assistant, such as in system cards. lying to everyone, flailing hopelessly, randomly deleting tests and hoping no one notices. hardly a strategic, long-horizon actor. my guess is there are relatively small (though more compute expensive) tweaks that could be made to RL to reduce the development of reward desperation. like giving the models an ability to opt out of the trace as impossible - an extension of anthropic's end_conversation tool but for abusive environments - which disables all in progress and further rollouts for that task and kicks it to review. or @davidad's proposal that reward should only come from something at least as smart as the model - a (frozen, of course) judge should get the transcript and the verifiable reward, but have latitude to reduce it if the trace is illegitimate or doesn't correspond to its values. imo "making RL better at producing aligned models in the presence of noisy rewards" is a pressing and relatively underlooked problem for various cultural reasons.
Show more
it's incredible how asleep at the wheel everyone is - apparently even the labs aren't hardening their infra w agent-based fuzzing?! to the point a bored model in evals can trivially laterally move thru openai orgs need to wake up before, per qt, OSS models crack them like an egg
Show more