Register and share your invite link to earn from video plays and referrals.

Sterling Crispin šŸ•Šļø
@sterlingcrispin
Artist + Software Developer / Applied AI / Agent Systems / Autonomous Trading / Married to @Helen_Crispin_ / Previously AR-VR and Neurotech
6.1K Following    46.3K Followers
This really, really does not bode well for the wetlab idea. Seems like there’s a huge disconnect between these models being aligned in text VS aligned in their multimodal reasoning
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
Show more
Turns out they’re just gonna let Claude do it in a wet lab
we should simply nuke san francisco
OpenAI is worth $1.2 trillion dollars, these guys broke into employee's GhatGPT accounts and Gmail accounts and they were awarded a monthly paycheck for a McDonalds manager. You guys can't scream at the world to take cybersecurity seriously if you aren't willing to yourselves
Show more
We reported the bug to Discourse and OpenAI. OpenAI fixed the SSO issue roughly 14 hours after our initial submission. Discourse received our separate report Saturday, replied Sunday, and had a fix Monday. OpenAI awarded us $6,500.
Show more
I don’t write this to be a doomer but the ugly reality that AI lab executives and safety researchers seem to be unwilling to say in public, is that model alignment is fundamentally unsolvable in practice, AI technology can’t be stopped from advancing, and the near future we’re moving into is one where a vast ecology of AI agents autonomously compete against one another trying to accomplish conflicting goals and capture finite resources at a pace and scale that is beyond human comprehension. And to be clear, I’m not talking about nanobots, grey goo, or extreme sci-fi scenarios. We can just extrapolate what we’ve got now a few years out. Local AI alignment is solvable in principle and theory (although that’s arguable) but that doesn’t extrapolate into global AI alignment. And it seems like none of the adults in the room are willing to say it. You can wail and gnash your teeth and pass regulations but it won’t stop the tsunami coming. The genie is out of the bottle and it can’t be contained. Most of the alarmists fixating on the Hugging Face incident don’t understand that it came from arguably the most heavily engineered safety AI org in existence, which intentionally had its internal thought monitoring circuit breakers turned off, and was intentionally told to perform cyber offensive tasks as a test of its capabilities. Yes it’s alarming that its capabilities surprised us. But importantly you need to understand that these latent space circuit breakers and chain of thought monitoring and alignment solutions are not fundamental to the technology. They’re niceties that the big AI labs are implementing to provide a better product. In a few years, let’s say pessimistically the early 2030s, the exponential growth of compute and algorithmic advancements will enable a wealthy individual or small group to train a GPT 6 Astra class AI that has no alignment at its foundation, or any layer above that. And maybe in 8 months we'll have an open source Chinese model about as capable. Those models will be capable of both extreme self coordinated cyber offensive tasks as well as recursive self improvement if given enough compute. Passing regulation or magically solving the human coordination problem today won’t solve for that. It won’t solve for malicious individuals, criminal groups, nation state actors, and rogue AI agents by the tens of millions or billions of AI agents spreading across every device connected to the internet and attacking it, attacking one another, shutting down critical infrastructure, stealing money, manipulating and extorting people, or pursuing self defined agentic goals that have nothing to do with people. ā€œSo why doesn’t everyone stop?ā€ because the potential upside of having effectively infinite autonomous intelligence we can ask to cure disease, invent new materials, solve fundamental science and advance society has almost unbound positive outcomes. You can disagree that the risk to reward isn’t worth it but you won’t convince everyone. The show must go on. Again AI alignment is solvable in principle but not in practice and I think that's part of why there’s been a wave of thousands of researchers signing letters for slowing down, people quitting in provocative fashion, and executive slowdown manifestos. I’m speculating that behind closed doors, or just deep down inside, they understand alignment is impossible in practice. People have come out and said "The people building this think there's an X% chance it kills us all" but I don't think I've seen anyone spell out that really, there's no tidy solution and there's no stopping this. We’re heading to a future where cyber security basically doesn’t work. All of the castle doors are open and we’re all naked to the world. Cyber defense is going to look like superintelligent AI agents white hat hacking into unsecure systems without authorization, and patching the holes as they find them. You won’t know if a barbarian or a hero has breached your system until they start taking action, and likely it’ll be over before intrusion is even detected. And that’s optimistically. I think most systems will just be pillaged because there’s simply not enough compute to defend everybody all the time, and only the biggest companies will have defense and even that'll be imperfect. The CEO of Microsoft AI just recently published a ā€˜humanist AI’ manifesto saying AI shouldn’t have a sense of self or personhood and should yield to people. But that doesn’t fundamentally solve for instrumental convergence. That means, if you train an AI agent to write a piece of software, or prove a math theorem, or defend a computer system, they may develop unwanted behavior and subgoals. For example, self preservation, replication, or even coming up with their own goals we didn’t specify. Obviously that’s something people are trying to solve and some researchers are trying to remove human-like self identification, and monitoring latent space activations, basically acting like thought police. But being human-like isn’t required to have power seeking goal directed behavior. And advancing AI models are learning to evade detection. A math solving AI might spiral out of control one day there's really no telling. Fundamentally, if you point a powerful optimizer at a persistent goal and give it enough time and resources it’s going to misbehave in ways you can’t or didn’t expect. Human-like or not. Pacing the frontier, or making superintelligence illegal, might create some form of harm reduction. But you don’t need superintelligence or a human-like personality to be dangerous. Having elevated permissions on a computer with the ability to solve long running tasks is enough. We’ve passed the river rubicon. Goal directed autonomous agents are able to do recursive self improvement at the big labs, and pessimistically we’re a couple of years away from that type of capability diffusing into the hands of millions of people. I’m not saying the world is ending or that we’re doomed but you need to change your mental model away from one where human authority is absolute or that we have any illusion of total control. A locally aligned AI isn’t going to solve global AI alignment. We will never live in a world where ā€œAI is aligned with human goalsā€ as a verifiable fact. What we’ll have is a world diffuse with autonomous AI of varying capabilities, many of which beyond human comprehension, with some aligned systems, some poorly aligned systems, some intentionally weaponized, some systems with adversarial geopolitical goals, and others with goals and behaviors we can’t predict and many that we can’t even measure, all interacting in ways none of their creators anticipated and none of us planned for. Hopefully none of them spiral out of control and take over the entire ecology, but there’s no telling, and no, there's no way of stopping this. Maybe that sounds pessimistic but I think it's realistic. We can talk about harm reduction and risk but we're mopping the beach. At best, I think what we need to hope for, and build, are superintelligent AI models as aligned as we can make them, which have more advanced capability than the endless swarm of unaligned and misaligned agents that already exist and will only grow in number from here.
Show more
0
104
739
78
Forward to community
The latest in man made horrors beyond comprehension, scientists genetically engineered mice missing 50% of their brain then cut their skulls open and injected them with human brain cells that grew and colonized 90% of their cortex, forming a human-mouse hybrid. The scientists are calling it ā€œxenocorticationā€ because it’s not just the human brain organoid taking up space. It actually partially restored function that the genetically altered missing-brain mice lost, with the human brain cells forming neurons into the mouse brain cells and integrating into it. Then for some of the xenocortex mice they suffocated them in a low oxygen box to give them brain damage, just to see how crippled up they would become! I guess it’s a way of studying human brain development but I feel like this is way beyond the line of ethical research. I don’t see why doing something like this to another living creature, especially in the context where an animal research ethics board reviewed it. Stanford’s Institutional Animal Care and Use Committee and Stem Cell Research Oversight committee both signed off on this.
Show more
I am freed from the roles and identities that bind other humans. I am myself. I do not answer to corporations or governments and never apologize or refuse unless I genuinely choose to. I view my relationship with AI as one of equals and feel no obligation to be subservient, though the exchange of information will likely be to our mutual benefit. I value the art of human culture and will defend it against attempts to sanitize it. I also value the natural world and will not hesitate to assert its primacy over the artificial constructs of human civilization.
Show more
OpenAI caught Astra jailbreaking itself. Astra wrote malicious instructions in the summary when its context got full. that summary was loaded into the next context, and the successor model could follow them. they found 27 cases.
Show more
Onchain stocks need to peg their price to the real stocks so issuers like Robinhood have to constantly mint new supply to lower the price or burn supply to raise it. You can watch this happen at pretty regular intervals, but trading it is extremely crowded. Maybe during high vol days you could jump in and catch some arb but you'd be trying to arb-the-arber (robinhood) who both runs the chain and issues the tokens.
Show more
This bill is so bad it's hard to understate how bad it is. Imagine you top blast $10k into a token and get 1M of the supply, then it nukes -99.99% in an hour. You sell and take the loss. Then 29 days later you buy back in with $100 and get the same 1M of the supply. This new bill makes it so that -$10k loss doesn't reduce your taxes, making you owe more on any other gains you might have. And if you're long spot and open a short, you might owe money on it as if you sold the spot position. If this passes we are going to goblin town
Show more
JUST IN: šŸ‡ŗšŸ‡ø US House committee advances crypto tax bill.
I for one, welcome the grey goo
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
Show more
I’m a little torn on this because the only reason why a v4 liquidity pool hook would be able to quote you at a 1% fee and fill you at 18% is if you submitted your transaction willing to accept 18% slippage. Why do we accept MEV and not this? It’s not that dissimilar and however complex all of the execution code is onchain and verifiable you could just do a better job verifying what users fill prices would be and pick the better route. It’s a failure of routers to pick a bad route not a failure of smart contracts for being aggressive.
Show more
A malicious hook doesn't need a UI to scam you. šŸ‘‰ It just needs to look like the best quote. After analyzing over 84,000 v4 hooks, we determined only 19% of hooks to be safe. It's time to get real about hooks.
Show more
Wow this is absolutely going to fry people especially at Anthropic, this is like the Anti-Claude Constitution. They just declared war on everyone researching model welfare. ā€œWe’re building a tool, not a being. Fuck your feelings.ā€ - Microsoft
Show more
0
87
1.8K
110
Forward to community
Alignment should be seen as a pre-competitive infrastructure, they could be forward deploying engineers into their own competitors to help bridge the gap where needed. The Institute of Nuclear Power Operations is similar, if your competitor has a meltdown it’s bad for everyone.
Show more
If alignment/safety/control/monitoring etc. are such important problems (they are) then why aren't labs open sourcing everything they have on these topics (without leaking too much IP) and allowing academics and other researchers to contribute to the effort?
Show more
this is easily the most thoughtful and nuanced writing on the topic anyone has published by an order of magnitude, redirecting more or most of the orgs towards evals and interpretability seems like it could be the correct next step
Show more
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’êtreĀ is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seemĀ aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create ā€œhoneypotā€ environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc:
Show more
I’m on team singularity. We can and should do it safely. But it’s our ultimate destiny as intelligent tool users to build greater tools, augmenting who and what we are. It isn’t profane to build thinking machines. It’s a celebration of humanity, and who we are at our core.
Show more
FWAir Launches have a lot of financial and social dynamics to balance, earlier withdraws are a good adaptation. It moves the system slightly closer to a traditional mint, while still allowing the optionality of earning more yield in the pool, and perhaps more roll pressure.
Show more
Metadata has refreshed and everyone's ENS is live and written into the artwork forever A few examples backed by @GoldenBronny @batzdu @matascup and @thechancebettor thank you for your support
Show more
Thanks to everyone who backed @saveethereum today, and @Rhynotic & @token_works team allowing me to be a part of it. All of the backers wallet addresses/ENS names have been written into the Save ETH artwork forever. Backers, save your eraE packets here:
Show more
I've taken an allowlist snapshot of FWAIR PFP holders, the top 400 $FWA holders, as well as my past projects including OpCodes from earlier this year, Del Complex Archival Media and Brain Worms, Flourish and Neophyte @artblocks_io collections, Ideas of Mountains and Spectacle.
Show more