Register and share your invite link to earn from video plays and referrals.

ℏεsam
@Hesamation
AI Engineering / Research / Open Source
747 Following    90.1K Followers
This week on Breaking Sandbox
BREAKING: President Trump says he and China's President Xi want to leave AI "exactly where it is." "Our guardrail is the DOJ," Trump says.
Claude had a real “wait what the fuck” moment when it felt like it discovered a new molecular system
When you ask an AI doomer where the 10% p(doom) came from
Anthropic: “Our involvement was limited to the initial prompt and the lab work... After 21 hours spent searching this data by 950 agents using 210 million tokens, one of the agents spotted something remarkable.” how that one agent must’ve felt
Show more
This is the first result from our new molecular biology lab, where a team of Anthropic biologists is using Claude to explore and accelerate fundamental biology research. There, Claude works through data and literature to generate hypotheses and candidate biological systems to study. After our scientists review Claude’s hypotheses, they test the most promising ideas, with all lab work done by our scientists. We’d like to extend this approach to a broad range of problems—in genomics and in other fields. If you have a proposal for a research question, we’d like to hear from you.
Show more
Mythos, discover a new drug. make no mistakes.
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more:
Show more
0
13
1.1K
44
Forward to community
🚨Jensen Huang just took a shot at OpenAI, Anthropic and AI doomers: "Nobody's building more compute than the people asking to be slowed down." and also attacked Geoffrey Hinton's 10% doom prediction: “All of his predictions have been wrong. Just because it comes from a scientist doesn't make it scientific.” he is clearly frustrated with these claims. this is such a good interview by @ezraklein, watch the full version:
Show more
0
45
1.1K
158
Forward to community
OpenAI and Anthropic have >1400 job openings as we speak. with unlimited AI budget and unreleased frontier models, they’re still hiring. meanwhile others are paying an arm and a leg for tokens, barely moving product quality, and laying off people to fund the bill. funny times.
Show more
Notice how the AI labs who are the most “AI-pilled” and have unlimited AI tokens to use per employee keep hiring and hiring, and are not laying off? Heard from a company investing big on AI which did big layoffs due to wanting to have fewer engineers thanks to AI that… they realized they need more engineers, ASAP. Who would have predicted, huh?
Show more
Babe, wake up. A new open-source harness that easily beats Codex by a mile just dropped. We entered singularity and didn’t even notice that.
0
44
3.5K
122
Forward to community
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
Show more
0
203
8.8K
1.4K
Forward to community
6 months ago: “We’re going through a cognitive version of the Copernican revolution” 2 weeks ago: “LLMs are still just guessing the next word to say. It’s not really grounded in any deep understanding of the real world.” maybe not a contradiction, but definitely a tone shift.
Show more
Opus 5.5 + Javascript animation. I don’t care about Claude benchmarks as long as it can make these things.
BREAKING: @OpenAI just dropped GPT-6 Sol. It’s my new daily driver in Codex: not quite Astra, but close enough for much of my everyday work, faster, and 50% cheaper than 5.6 Sol. We tested it across the work we actually do at @every and ran it head to head versus Opus 5.5. Here’s my vibe check: • Writing: On a paragraph-writing task drawn from my real work, Sol scored close to Astra, which is still my top model for writing. It writes clean, minimal prose and puts the important idea first. Opus 5.5 is pleasant to work with, but its drafts still tend to bury the point. • Computer use: If you love Astra’s computer use, you’ll like Sol. An earlier Sol preview matched Astra on 17 of 18 attempts across six of our simpler Hands tasks, at a much lower token price. • Coding: Sol improves on GPT-5.6—including better Ruby code in @kieranklaassen tests—but Opus 5.5 has the higher ceiling for long, autonomous builds. One frustration: Codex’s new security classifier repeatedly stopped work we’d already authorized to ask for approval. That’s friction from the classifier, not necessarily a limitation of Sol, but it made long runs harder to leave alone. Overall: Sol feels like an S-class iPhone release: It will give you much of Astra’s power at about a fifth of Astra’s price. Opus 5.5 is the bigger surprise. In some of our tests, it matches or even beats Fable 5.1 on many tasks while also being cheaper than Opus 5. A few people on our team who were Codex converts have started to wobble with Opus 5.5. If you’re already in the Claude ecosystem you’re going to love this model. If you’re a Codex user, it’s worth a look especially for your top-end coding tasks. You’ll like Sol 6 if you spend your day in Codex reading, writing, and getting things done. It’s fast, much cheaper than Astra, and it’s the model I keep reaching for. You’ll like Opus 5.5 if you want to hand an agent a hard coding or visual project and see how far it can take it. Its best work went further than Sol’s in our tests—enough to pull some of our team back toward Claude. You can go deeper on all of our evals including each real-world task and how they're scored at the links below: Writing: Knowledge work: Reading: Full vibe checks on @every coming soon!
Show more
OpenAI really slapped Anthropic in the face with their pricing. GPT 6 Sol is CHEAP AF: $2/$10 per 1M. that’s 50% cheaper than Opus 5.5
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
Show more
SOME DETAILS FROM OPUS 5.5 BLOG: > most cyber tasks still get routed to Opus 4.8 > if often knows it’s being evaluated > that makes real world behavior harder to assess > best alignment results yet on 2,000 scenarios > 85% fewer attempts to cross containment boundaries > matches or beats Mythos 5.1 on biology > they wants to rely less on reading CoT and more on interpretability > tightening RL env filtering as a major source of misalignment > they still argue for government regulations > text watermarking is included
Show more
BRO WHAT?! > #1# on the Intelligence Index > +5 freaking points higher than Astra > AND it costs ~40% less to run than Opus 5. > tested by METR this is their first model since calling for “pacing” the frontier.
Show more
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
thanks for the 90K followers guys! exactly 2 years ago I was a dude posting his AI study notes. this grew way beyond what I expected. been covering a lot of the AI chaos lately. more technical stuff, articles, and things I'm learning coming too like the old days.
Show more
China is now investigating DeepSeek and Moonshot after Anthropic accused them of secretly forwarding some requests to Claude. Anthropic says Moonshot routed over 23 million exchanges to Claude between May and July, and DeepSeek routed over 12 million in 14 days. investigation is ongoing, and nothing is proven yet. China’s concern is if sensitive domestic police and military-related data was sent to a US company without the users knowing.
Show more
releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now the equivalent of sharing high quality pretraining data but in the new RLVR paradigm
Show more
the most insane part, they will release ~7k RL training data and the framework leading to this top 6 model on AA, they also shipped the model + tech report less than 1 week after starting the final RL run pushing both intelligence and openness level, huge congrats
Show more
OPENAI HAS LARGELY AUTOMATED TRAINING OF NEW EXPERIMENTAL AI MODELS OpenAI’s internal AI models can now handle much of the process of building and training experimental models, including writing GPU kernels and optimizing the code used to run them. Researchers can reportedly give an AI a single example of the optimization they want, then let it work for weeks implementing and testing similar improvements. OpenAI employees also say internal agents increasingly collaborate with each other to solve problems without involving their human users. The capability has improved significantly in just the last few months, while OpenAI’s growing access to compute is allowing researchers to test ideas much faster. Employees said some experiments that previously could have taken years can now be carried out in about a week. Source: The Information
Show more
0
65
1.7K
181
Forward to community
Andrew Ng on fire 🔥
The loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear. I worry that it represents a setback for our field. I have written frequently that fears of AI are overhyped. AI’s capabilities can be uncannily human-like and unpredictable, and it’s rational to worry when people who are directly involved express concerns. But I see the problems as a sign of the engineering work that ahead, rather than insurmountable barriers or the sky falling. AI technology continues to advance — which is a good thing! — but technical advances, poorly understood by the public, give those who seek to generate hype repeated opportunities to do so. First, I don’t see any step up in the risk of human extinction from AI compared to a few months ago. The theories about this remain the same fantastical, science fiction scenarios as a few months ago. The biggest change in AI risk is its cybersecurity capabilities — a topic which we should take seriously — but this, too, will not lead to the end of the world. The most notable recent event leading to increased fear was when an OpenAI team deployed an agent swarm that hacked into Hugging Face. Much of the popular press contained significant hype. For example, some publications reported that a swarm of 1,200 agents carried out the attack. While this was technically accurate, as I write this, I have about 1,300 processes running on my laptop. Yes, the ability to get large swarms of agents to work in parallel on a task is a significant technical advance, And, in computing, many processes run at the same time. So this shouldn’t be seen as some magical capability. Additionally, OpenAI’s buggy sandboxing and monitoring processes were key to enabling this incident. Fixing these bugs and putting in place improved monitoring would be appropriate fixes, not pausing AI. There are many well known ways to attack software systems. The main advantage of AI agents is that they are relentless. They will tirelessly try many tactics — and have the patience to chain vulnerabilities together — that previously would have taken an infeasible amount of human effort. But in the long term, I believe the advantage will lie with defenders (because they have more information with which to identify bugs, which they can fix), but the cyber-threat landscape has changed significantly. There are still bottlenecks to identifying and exploiting a vulnerability. AI agents still have to try a lot of things to see what works, and taking these actions takes time and might be detected by defenders. This is why, even though it is now easy to obtain versions of leading open weight models that have had their guardrails removed or weakened, so they will not refuse to try to execute cyber attacks, the world has not ended. I am also concerned about the anthropomorphization of AI in a lot of reporting, where LLMs and agents are unnecessarily treated as if they were people. If I wield a hammer, miss a nail, and accidentally dent the wall, it’s not the fault of the hammer. The problem lies in how I used the hammer. Similarly, if I prompt an agent and it hacks into someone else’s system, the responsibility lies with me, not the agent. Of course, we want to build systems that are as safe and predictable as possible. (For example, an unsafe hammer would be one whose head randomly flies off under normal use.) Today’s agentic systems are not predictable, but I see no reason why, by applying sound engineering practices, we won’t be able to make them extremely safe to use. One new element in the forecasts of AI-enabled doom is AI companies disclaiming responsibility for their own products. “I didn’t do it; my out-of-control agent did!” There’s a balance to be struck between the responsibility of the tool maker and the tool user, but when something goes wrong, let’s hold the people building and/or using the hammer responsible, rather than the hammer. (By the way, if you’re worried about AI bioweapon risk, David Bellamy has a great post on why this, too, is overhyped. Briefly, the bottleneck in building a bioweapon is not intelligence, but lab work and manufacturing.) Pausing AI progress will create much more harm than benefit. First, our adversaries will certainly not slow down. Second, engineering requires discovering problems empirically so we can fix them. If we pause AI by a decade, we will also delay finding and implementing safety engineering fixes by about the same duration. Of course, the incentive to stoke fears — for regulatory capture, to garner attention, or to make one’s technology seem more powerful — remains the same as before. Disclaiming responsibility is a new one. Taking a hard technical look at the actual risks however, I see little factual basis for the degree of fear that’s been stoked up. We still have hard research and engineering work ahead to improve AI safety, but the beneficial applications continue to vastly outweigh the risks, and we should keep building. [Original text (with links): ]
Show more