Register and share your invite link to earn from video plays and referrals.

Lance Martin
@RLanceMartin
476 Following    39.7K Followers
Opus 5.5 is out! excellent at coding, great at writing, reduced pricing. run '/claude-api migrate' in the latest Claude Code to update your application code. run '/claude-api prompt-audit' to ensure your skills + prompts are well-tuned. short video overview:
Show more
i recently added this command to the claude-api skill. run it in Claude Code to fix common prompting "anti-patterns" that can hobble frontier models: /claude-api prompt-audit patterns include: 1/Verification rituals. Instructions like "double-check your work” or "verify twice before responding” are often taken literally by frontier models and can waste tokens. 2/ Thoroughness and emphasis boosters. "Be maximally thorough," "CRITICAL: YOU MUST ALWAYS…" can lead to verbosity and extra tool calls when working with frontier models. 3/ Mandatory procedures and scratchpad scaffolds. Fixed step processes (e.g., "think step by step in a scratchpad") or reasoning templates are rituals that frontier models don't need. This scaffolding can stack on top of native reasoning and use unnecessary tokens. 4/ Stale examples. Few-shot examples tuned to an older model's failure modes can teach a frontier model to imitate long reasoning chains on requests that don't need them. 5/ Contradictory rules. Frontier models are better at instruction following. Contradictory instructions ("always refund within policy" vs. "never issue refunds without escalation") can be followed more literally by frontier models, resulting in degraded performance. 6/ Dated configuration. Settings written for an older Claude generation (e.g., manual thinking budgets) can be rejected by the Claude Platform with newer models. these patterns accumulate in prompts over time, and can quietly degrade performance when upgrading to newer models. a common reason is the frontier models are better at instruction following, so these anti-patterns steer them to spend unnecessary tokens. example: i tested a migration from Opus 4.8 to Opus 5 on an internal customer support benchmark. with Opus 5 (and other frontier models like Fable 5.1), verification rituals ("verify twice") use unnecessary tokens by duplicating work. emphasis boosters ("be maximally thorough") become dozens of unneeded searches. applying prompt audits can improve performance and reduce cost (as shown in example attached and will be sharing a full write-up soon). also, the skill is also open source and some of this guidance likely applies generally across frontier models
Show more
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work. We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets. We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
Show more
0
607
12.5K
1.4K
Forward to community
This post is a good example of how Dario responds to people in slack - candor and substance. Recommended reading!
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers.  Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up.  This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).  Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).  By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”.  The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure.  I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity.  This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
Show more
0
69
1.2K
40
Forward to community
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers.  Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up.  This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).  Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).  By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”.  The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure.  I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity.  This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
Show more
0
1.2K
9.3K
912
Forward to community
As we build our agents, we are all trying to get the most intelligence per dollar. Check out these tips we just launched for Claude Platform. Oh, and Claude created a fun video in one shot!
Show more
This is part of working with the EU AI Act, other labs are adding similar watermarking. It’s hard to identify AI-generated text, and this gives people better tools to do that. We’ll also be a shipping text detection API that you can use yourself.
Show more
0
563
1.2K
55
Forward to community
MCP 2026-07-28 is live and it's the largest update to the protocol since launch. MCP is now stateless, making it easier to deploy and scale remote servers.
0
367
13.3K
1.4K
Forward to community
New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are used to keep data private. Read more:
Show more
0
452
6.3K
839
Forward to community
There’s been a lot of speculation about where we stand on open-weights models. We’ve outlined our views in full here:
0
2.8K
8.2K
1.1K
Forward to community
lol this talk is from AIE 2024 (before I was at Anthropic). im glad graphs are cool (again)!
Anthropic engineer just released a 2-hour workshop on "Graph Engineering" for agentic systems: “80% of our engineers are using self-improving loops. Now everyone is building agentic graphs.” • 00:00 - Introduction to RAG & Graphs • 06:39 - Core of "Graph Engineering" (state, nodes) • 14:29 - 3 feedback loops of Graph agents • 23:06 - Agent evaluation with Graphs • 36:29 - Agent cycles in graphs • 1:15:22 - Agentic RAG & agent context • 1:41:20 - Evaluation datasets based on Graphs This 2-hour workshop will replace 10 paid courses on agentic engineering. Watch it today, then learn how to become a Graph Engineer in the article below.
Show more
Big news from our internal writing benchmark (early results): Claude Opus 5 by @AnthropicAI is now #1# for writing in our editorial voice, at 2817 Elo, surpassing Claude Fable 5 and Kimi. Already! That is a jump from #15# to #1# over its predecessor (if we take all thinking variants into account), Opus 4.8, at the same API price. Reasoning effort actually matters this time. At default effort it lands #6#. At max effort it takes the top spot, thinking for over three minutes per script. Seems obvious, but it wasn’t the case for 4.8, though it is for Fable.
Show more
We removed ~80% of the Claude Code system prompt for our newest models, this is what we've learned about writing system prompts, skills and Claude.MDs for them.
0
463
16.1K
1.9K
Forward to community
Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed novel behavior that allows Opus 5 to solve previously unbeaten environments, outperforming Fable
Show more
Opus 5 is a great model for async / long-horizon work. Check out my talk from @aiDotEngineer on a few useful patterns : 1) split the brain + hands, 2) self-correction loops, 3) memory with dreaming, 4) org-level agent harnesses.
Show more
Props to OpenAI for publishing this post on some safety and alignment issues observed in internal deployments - there are many counter-incentives to publishing stuff like this, but by making it public we all get better info about safety at the frontier.
Show more
We're offering grants of up to $50,000 in Claude usage credits to researchers accelerating cures for rare diseases. This is our first focused call within AI for Science, our program supporting scientists using Claude to speed up discovery.
Show more
0
766
8.5K
1K
Forward to community
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Show more
0
1.7K
44K
5.4K
Forward to community
working more on a post about what we learned doing this and how you can apply this to your skills and system prompts
0
55
1.2K
53
Forward to community
Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to predict, which is why we rolled it out to subscription plans in stages, extending access several times as we secured additional capacity.
Show more
0
5.3K
53.6K
6.2K
Forward to community