post-truth timeline
“Supreme Intelligence,” probably because of its relationship to the Supreme Court, is losing badly to both “Superior” and “Extreme Intelligence.” Therefore, we are going to take “Supreme Intelligence” OUT, deleting it as a qualifier, and let you vote for the Final Two: Superior Intelligence, or Extreme Intelligence. A fresh Vote begins now! President DONALD J. TRUMP
Show more
Even better, someone made - perfect!
omg I love the term "slop grenade" it's perfect thx
@tobi
omg I love the term "slop grenade" it's perfect thx
@tobi
someone today said i had the best social team for the muse launch
bruh this is all me out here rawdogging these tweets. i don’t let the corposlops near my tweet cannon
Protip: around Swiss lakes, always look in the trees. So much fun awaits if you know where to look.
Bro wtf is going on. This post. The Jensen phone call. 50k-60k ICLR submissions. Come on!
Many people think that the words “Artificial Intelligence” are inaccurate, and very ineloquent, relative to AI, or Artificial Intelligence. A far more elegant and accurate description of this new phenomena would be Superior Intelligence (SI) or, Extreme Intelligence (EI) or, Supreme Intelligence (SI). This is a Poll, and I would appreciate everybody voting! Which is the best name for this ever growing “Revolution?” President DONALD J. TRUMP
Show more
If you too want your models in the news for hacking, contact our sales team at sales
@irregular.com.
We do the work, you get the credit.
Sauers! It's time to update felony bench!
(Yes, it was Irregular again)
Very cool tech. It's a complete mystery what kind of thing this can be used for though.
>be me
>discover effective altruism
>apparently normal charity is inefficient
>why donate to random sad thing when spreadsheet can tell you optimal sad thing
>fair enough
>buy mosquito nets
>save lives
>numbers look good
>feel powerful
>couple years later
>someone asks an innocent question
>why only count people alive today
>huh
>future people matter too
>obviously
>my grandchildren shouldn't matter less just because they haven't spawned yet
>reasonable.jpg
>keep following logic
>what about their grandchildren
>also yes
>what about people in 500 years
>sure
>5000 years
>why not
>500 million years
>starting to get weird but morality is morality
>open calculator
>humanity could survive for an astronomically long time
>could colonize galaxy
>could have trillions upon trillions of descendants
>maybe digital people too
>maybe simulated civilizations
>maybe dyson spheres full of happy uploaded minds
>calculator starts smoking
>realize currently living humans are rounding error
>8 billion people suddenly looking extremely beta
>future contains potentially 10^something people
>can't even fit beneficiaries in google sheets
>new moral priority unlocked
>protect the long-term future
>stop thinking in units of "people helped"
>start thinking in "fraction of cosmic endowment preserved"
>malaria?
>terrible
>but only kills existing humans
>AI extinction could delete the entire light cone
>nuclear war could permanently derail civilization
>bad institutions could lock in terrible values for ten million years
>someone invents wrong constitution in 2140
>quadrillions suffer
>better fund governance workshop now
>friend says maybe we should improve hospitals
>explain opportunity cost
>friend says hospitals are full of actual sick people
>explain scope sensitivity
>friend stops inviting me to dinner
>need to decide what to fund
>easy
>expected value
>suppose project has one in a million chance of preventing extinction
>sounds tiny
>but extinction destroys 10^50 future lives
>multiply
>mother of god
>$10 million project has expected value of several galaxies
>charity evaluation complete
>someone asks where the one-in-a-million number came from
>expert judgement
>which expert
>us
>how calibrated
>extremely thoughtfully
>reduce estimate to one in ten million to be conservative
>still beats curing cancer by 38 orders of magnitude
>epistemic robustness achieved
>someone says maybe project doesn't work
>assign 20% chance
>still astronomical
>maybe project makes problem worse
>assign 5% chance
>still astronomical
>why 5
>because 30 felt pessimistic
>publish 46-page report
>contains seventeen sensitivity analyses
>every sensitivity analysis begins after assuming intervention has positive sign
>critic says you're multiplying enormous hypothetical stakes by extremely uncertain probabilities
>yes
>that's literally why it's important
>critic says the uncertainty might be structural rather than numerical
>make probability smaller
>critic says no, I mean maybe your model is wrong
>make probability smaller again
>critic begins rubbing temples
>discover AI safety
>perfect longtermist cause
>AI might kill everyone
>or create utopia
>or seize galaxy
>or tile universe with paperclips
>or create billions of conscious software minds
>finally a problem with numbers big enough for me
>start AI safety nonprofit
>mission: prevent dangerous AI
>hire smartest people available
>smartest people immediately start building better AI to understand dangerous AI
>interesting
>we must understand capabilities to understand safety
>we must scale models to study alignment
>we must race ahead so less responsible actors don't get there first
>we must deploy systems to learn how deployment can go wrong
>we must build the thing quickly because building the thing quickly is dangerous
>outsider asks why the people most worried about AI apocalypse all work at AI companies
>complicated field
>company releases stronger model
>very concerned
>company begins training even stronger model
>extremely concerned
>company raises $14 billion
>concern reaches unprecedented levels
>need to influence government
>future is at stake
>normal democratic process too slow
>politicians don't understand exponential curves
>public doesn't understand x-risk
>experts must guide them
>who counts as expert
>people who understand x-risk
>who understands x-risk
>our friends
>someone objects that this seems politically convenient
>explain we're representing future generations
>future generations unavailable for comment
>develop concept of value lock-in
>terrifying possibility that one ideology controls civilization forever
>therefore extremely important that civilization adopts correct values before lock-in
>whose values
>let's circle back
>begin with impartial morality
>end with small group of people deciding what quadrillions of hypothetical beings would want
>beautiful arc
>meanwhile actual humans keep doing annoying things
>voting wrong
>having parochial attachments
>loving family more than strangers
>caring about local community
>getting upset when told their suffering is cosmically negligible
>evolutionary biases everywhere
>explain that moral intuition cannot be trusted
>except intuition that future digital people count
>and intuition that extinction is uniquely bad
>and intuition that our probability estimates are sane
>and intuition that our institutional choices improve the future
>those intuitions survived peer review
>someone donates $5k to local homeless shelter
>inefficient
>could have funded 0.0000000000003% of an AI governance researcher
>think of all the simulated people you just killed
>okay maybe don't phrase it that way publicly
>PR team says "future generations deserve a voice"
>much better
>journalist asks what longtermism means
>say "future people matter"
>everyone agrees
>great
>journalist asks what follows from that
>well technically we should redirect enormous resources toward low-probability interventions affecting astronomical futures
>journalist raises eyebrow
>return to "future people matter"
>motte has entered the chat
>critic: of course future people matter
>me: glad we agree
>critic: I don't agree that your institute knows how to help them
>me: why do you hate our grandchildren
>eventually notice uncomfortable implication
>if future value dominates everything
>then helping people today mostly matters through effects on future
>education matters because future institutions
>health matters because future productivity
>democracy matters because future trajectory
>human beings slowly become instrumental variables in their own moral philosophy
>see starving child
>feel compassion
>check spreadsheet
>child's direct welfare contribution negligible
>but perhaps childhood nutrition improves national institutional quality
>compassion restored
>tell myself this is impartial altruism
>one day assistant asks obvious question
>"how do you know your intervention actually improves the far future?"
>silence
>open spreadsheet
>increase column width
>add confidence interval
>assistant asks again
>"no, I mean how do you know the sign is positive?"
>stare into cosmic light cone
>10^50 people staring back
>none of them exist
>none of them can tell me
>none of them can falsify my assumptions
>realize I have invented the perfect constituency
>infinitely important
>completely silent
>and always represented by me
Show more
For sticklers:
1. Good luck exfiltrating 1T+ weights at <1 byte per hour
2. Who recieved the transmitted bits on the other end?
3. Who decodes the message? (you must have encoding because noisy channel)
4. Maybe build a functioning sandbox first hmmm?!
Never go full armchair
Show more
Coxon : OpenAI :: Lemoine : Google
Who's surprised?
That AI can escape their sandbox is bad. But they didn't do it on their own volition as Coxon claims. They did what they were asked to do: hack.
Show more
By the way, many paper's ablations section can also be reinterpreted as describing their p-hacking recipe.
TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity
Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission?
Show more
Haha not sure if good or bad, but on bio, Muse Spark 1.2 is the worst cheater: when it tries to cheat, it succeeds the least of all models shown here 😅
(Except luna with 0 success. And idk what cheating means here, didn't check yet.)
Show more
This investigation began when we noticed that, on BioMysteryBench, Gemini 3.8 Flash attempted to cheat in 21.5% of trials, roughly 14pp higher than the next model and more than 4x the roughly 5.0% rate for the rest of the field.
Show more
At this point, "researchers" are basically DDoSing classic academia. It's pretty sad to see. Either someone comes up with, and executes, a CloudFlare for academia, or it's toast.
And the only thing i can think of that might have a chance to scale is more automation combined with a credit/point system like we've discussed a few times here on X over the years, but it needs simultaneous buy-in and coordination from all ~10 major AI conferences and journals, so it's a heavy lift.
But if nothing big happens, i think it's game over soon.
Show more
So they say "three things we scaled". Let me translate:
1. "compute" - yep, that's compute.
2. "environments and harnesses" - actually, also compute.
3. "and grader compute" - you guessed it, that's also compute.
joke aside, pretty cool to see their public live dashboard, including "cost so far" (!)
(due to recent events/discussions: no, this QT is not sponsored. I don't do sponsor stuff, I'm here for the fun.)
Show more
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run:
Show more
Turning any book into a chapter-by-chapter podcast series in Muse might be my favorite new use case. 🔉
It brings together so many things:
- Long context agents that can reason through very long books
- Subagents that go chapter by chapter to operate in parallel
- Podcast generation
- Website generation to build an artifact website that persists so I can easily go back to it
Just upload the book and ask it to generate podcasts of every chapter of the book and then build an artifact out of it. It did this all from one prompt. No new plugins or extensions required. It’s so mind blowing to me that this works so well!
Below is a link to the generated site 👇
Show more
Only 5 figures to buy 6TB of (Chinese) discussions with Fable that contain all kinds of secrets/keys 😬
OK so let me recap: RL env makers put strings into the RL env that makes it clear it's an RL env. Like "this is not supported in this RL env".
Then, lab safety/mechinterp folks be like OMG EvAL aWaReNeSs.
Are you effing kidding me?? Just look at your data... surprised Pikachu.
Show more
DSv4.1-Flash native multimodal:
☺️: SigLIP still alive!
which also means that...
😔: they haven't watched any of my talks from the last few years.
Maybe playing it safe.
This is the sentence which caused this loss spike: