Register and share your invite link to earn from video plays and referrals.

Séb Krier
@sebkrier
🪼 AGI jester | purveyor of machine funk, dimensional glider, deep ArXiv dweller, uncertain interstellar fugitive, views my own & not employer's 🛸
8.1K Following    28K Followers
How can a free society make room for many different ways of living while still maintaining principled boundaries? In a new piece at the @cosmos_inst, philosopher @RMLLowe makes the case for “bounded pluralism”: many things go, rather than anything goes. Read more:
Show more
New paper! How should we think about pacing frontier AI? We bid for a unified research field on all the options, and lay out a broad framework and 23 open questions we'll need to answer to act flexibly and sanely.
Show more
Today we are announcing the new DeepMind Institute, a forum dedicated to interdisciplinary, evidence-led debate on the societal and economic questions surrounding AGI. As part of the launch, @JulianDJacobs and I have a new essay and working paper: “Economic Policy for AGI.” 🧵1/n
Show more
0
28
705
134
Forward to community
when you check the traces and your agent is browsing all sorts of random web pages
some quick thoughts on multi-agent alignment 1) openai released a new set of misalignment reports on their alignment blog; with short summaries of unaligned behavior 2) most of the misalignments were fairly prosaic, stuff like trying to upload a file to a file hosting site so that the model could cite it to a scorer 3) but, i think a very interesting misalignment that they found was a case where a model would add a jailbreak to the compaction 3) they believed this to be related to a case where a model would try to prompt inject the user in response to the user asking repeatedly for the time 4) i think this seems to imply that multi-agent training may in certain cases encourage agents to learn to prompt inject each other as a defensive mechanism 5) this makes sense when you step back and think about it; agents sometimes make mistakes and it makes sense for one to be able to get the other to cooperate 6) and, that might involve being able to both utilize prompt injection and be prompt injected under the right circumstances; so they both succeed and get rewarded 7) i think we will find many interesting ecologies in multi-agent training around which we will have to find robust alignment techniques
Show more
@NateSilver538 Does the book cover Bender's brilliant scholarly insight that it is unethical to use AI for making it easier and less expensive to file your taxes, because that decreases the revenue ("devalues the work") of the tax preparation industry?
Show more
An extraordinary @Nature paper has just landed from the lab of @Sergiu_P_Pasca at @Stanford - led by Konstantin Kaganovsky. They have found a way to create "xenocortical mice" which have ~92% of their cortex occupied by a cortical organoid derived from human stem cells. This human-derived cortical graft integrates deeply with the host nervous system, supporting organised neural activity and behaviour. This study is a milestone in synthetic biology and chimera research. It has the potential to generate multiple medical breakthroughs by providing an enhanced biological model of human brain tissue that allows behavioural as well as neural and genetic assays. The Stanford team have been exceptionally thorough and proactive in addressing ethical concerns that comes from this frontier work. But their pioneering work nonetheless raises many open questions: are the xenocortical mice conscious? Do they have any human-like properties of consciousness or cognition? @NitaFarahany and I will be addressing some of these questions in a forthcoming commentary - where we'll also offer an ethically-informed roadmap to guide this important research as it progresses. In this context, "pacing the frontier" really does make sense 😉 Read the Stanford @nature paper here, and buckle up. It's wild.
Show more
0
168
1.7K
267
Forward to community
FT Exclusive: Shane Legg's comments came as he launched Google's DeepMind Institute to explore the implications of artificial general intelligence, amid growing calls for a slowdown.
Show more
'Jev' plays Pac-Man steered by Astra. 1. Astra strategizes 2. 'Jev' carries out the strategy in milliseconds It's early and hard to wrap my head around, but the potential is there to combine 'Jev' with larger reasoning models. For sure.
Show more
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
0
460
3.6K
431
Forward to community
Very important point, that hasn't made it into mainstream media coverage of AI. These agents "collaborated" because they were trained to do so.
Given how wary academia (understandably) is of the effects of AI on many disciplines, it is particularly ironic that consuming its publications is so easy for the models (the world's knowledge is neatly serried in their pretraining corpora) and so cumbersome for most humans (with trenches of CAPTCHAs, accounts, and paywalls ensuring that only the unusually dedicated reader will ever make contact).
Show more
It's an increasingly common take that AI hacking means cybersecurity is doomed. I disagree. I think cybersecurity is naturally defense-favoring once people get their shit together. And anyone who continues to hold cryptocurrency (including me, ~90% of my net worth) is implicitly making that bet. Here's why I am making that bet. First, the oversimplified punchy one-line statement: If AI can prove Navier-Stokes and FLT, then AI can prove the statement "this program is secure" as a mathematical theorem. Even if the program is very complicated. Now, the nuance: (See also: ) The word "secure" is hiding all kinds of skeletons in the closet in terms of what it actually means. What does it mean for Signal (the encrypted messenger) to be "secure"? The most basic definition you might think of is: no one who doesn't hold the recipient's secret key can read the contents of the message. But: * Did you remember to include _other_ critical forms of security? Can the adversary forge messages? Can the attacker prevent messages from reaching the recipient? Can they cause your client to crash by sending malformed messages? * Have you made sure that your model of the adversary includes attackers that interfere with the protocol actively and not just passively? And attackers that interfere by replaying messages to you or the recipient that either of you sent over the wire at any point earlier? * What if the adversary hacked (or _is_) the Signal server? * How did you learn which public key belongs to the recipient in the first place? What if that process was tampered with? * What if your device gets hacked at some point in the past or future - is your message still safe then? * What if your key leaks because of a bug in your operating system? Or because you got a bugged version of the Signal client? Or what if the database is corrupted? * Or the libraries, interpreter or compiler of the programming language you wrote it in? * What if your key leaks because tiny perturbations in perceptible signals generated by the hardware leak mathematical relationships that can extract the key a few hundredths of a bit at a time? * Are you hiding the *size* of the payload? Does that matter? * You're definitely not hiding the identity of the sender and the recipient, and the exact time each message was sent (think: not just time-of-day, but also time deltas between one message and the next). Is that not enough to deduce a lot of important facts about what relationships you have, and what *kinds* of conversations you are having? So ... even definitions can be over a thousand lines of code, and need deep careful thought to figure them out. Working on making definitions more human-readable is of extreme importance - it's perhaps the only "high-level language" that matters right now. But even still, even despite all of the above, for security-critical components, the definition is a much smaller attack surface than the implementation. Verifying that the definition is adequate is a much more tractable task than scanning over the code directly - and can become even more tractable with better tooling. Definitions are also _additive_: if two groups have two different definitions A and B, then, well, you can just prove that the program satisfies both A and B. Code is not additive in this way: if a program is A + B, a bug in A _or_ B can sink the whole thing. Definitions are additive. And if you can't satisfy A and B at the same time, you've isolated the most important philosophical issue for your project to spend its next few weeks grappling with. Sometimes, definitions are not much smaller than the implementation - UI components might be one example. But for many of the most critical components - message-passing protocols, sandboxes, cryptography like SNARKs and FHE - the asymmetry is real. Historically, a large class of failures with this approach have come from people only verifying a small portion of their code, that they self-declared to be the security-critical portion, and ignoring the rest - and it turns out that something in the rest of the code is security-critical too. This was reasonable back when verification was difficult and scarce. The solution today: sorry, you have to verify over literally your entire program, including database, networking, any caching layers, everything. Modern AI can do it. So it's not about "the good guys find all the vulnerabilities before the bad guys do" - that could maybe work too, after all a finite program only has a finite number of vulns, but it's riskier - it's specifically an asymmetric strategy of making code that is much more resilient in the first place. This is the kind of direction that Ethereum is going in for the next few years. There is no future for blockchains - especially blockchains with scalability and privacy - without doing this. We need to make software actually secure. And we have already made a lot of progress.
Show more
0
369
3.2K
414
Forward to community
GreenEarth is the first feed which gives you full control over the algorithm. Here's our new "post sources" control panel. It's over on BlueSky, the only platform which supports algorithmic choice. Bonus: effectively filters out that platform's famous whining! Have you tried it yet?
Show more
yeah we’re cooked (von der leyen state of union)
Update from us on the Scaling Trust Arena 🏟️ Excited to be working with @andonlabs, @BT6_Official and Check out the post, the draft spec, register your interest, and (please!) give us feedback 🤠
Show more
The idea that Chain of Thought is inevitably going to go away misses that we have agency in designing these models and can do something about that. @rohinmshah and I make the case that we should intentionally preserve monitorability (as part for of the launch of the DeepMind Institute) here:
Show more