How Anthropic's new results post would read without the PR:
Claude orchestrated open-source protein design models, PXDesign, RFdiffusion, Genie, BoltzGen, from a 30k-token expert prompt and 12,500 H100-hours of compute, and designed binders against 14 of 15 targets. Hit rates of 22–35% against a 10–15% baseline, where some of those tools already report similar numbers on their own.
The orchestration is genuinely impressive. But the open-source models did most of the lifting, and they came from the Baker lab, Columbia, MIT, ByteDance Seed, and most of them were already wet-lab validated before Claude touched them.
Which also sets the ceiling. All these generators share a single PDB-shaped training distribution, so calling four of them doesn't diversify away the blind spot, since they fail together. The targets that worked are the well-studied ones.
So the valid claim is that an agent can now drive this stack competently in the regime where the stack already works.
Instead, we got this announcement:
LLMs increase the complexity of codebases. They duplicate methods, write overdefensive code against impossible edge cases, & overoptimize too early. Can further training fix this? Naur's “Programming as Theory Building” says no. -- @pol_avec
1/
"Thousands of people around the world are working on AI outside of the dominant narrative, in ways that embrace their own values" 💯
"i returned to the field because I want to figure out how to use AI in ways that protect human creativity, autonomy, and problem-solving." 🙏
As the backlash against AI was growing stronger (and as my friends were becoming more fervently anti-AI), I decided to return to an increasingly hated field. Why? I agree there is a lot that is terrible about AI. 1/