Register and share your invite link to earn from video plays and referrals.

Kevin Simback šŸ·
@KSimback
COO @delphi_labs - building + investing in AI and crypto. Ex @IBM, @McKinsey, @CarnegieMellon, reformed CFA
925 Following    20.9K Followers
ā€œApplied AIā€ ≠ doing AI right It’s easy to apply AI the wrong way inside companies To apply AI correctly, it takes the right combination of leadership, business acumen, and technical expertise Most AI deployments lack 1 or more ingredients
Show more
I am done with this shit. It is over. The state of engineering right now is horrible. It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow? People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Humans in corporate are doing nothing on their own. Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude. There is no sense of victory. Nobody is resolving bugs. In reality, nobody is thinking anymore. Everything is done by LLMs. It is so soul-sucking. I would not mind it, to be honest, if we were at least given the time to check out the code and see what is going where. But no, the goal is to just ship. No matter what happens.
Show more
The two best tools I’ve found for this: 1. Orca @orca_build 2. Herdr @herdrdev I personally use Orca but both are great
I keep switching between Codex and Claude. Why hasn't someone built a harness that works with both?
You can’t tell me that we get several trillion $ AI companies but don’t get trillion $ robotics companies as well Also bulled up after reading the @citrini piece on my flight
No one is ready for the Robotics bull run.
I am incredibly bullish on AI education I have two kids in school but live in one of the worst areas for public education I am grateful to have the means to send them to private school, but it’s not enough So I’ve been augmenting at home with AI, a ā€œvibe coded Alpha Schoolā€ of sorts And the tools are getting better, I test as many of them as I can, and I’m excited to try this one out Because I wake up every day thinking about how I can give my kids the best chance of escaping the permanent underclass Which means teaching them to think independently while leveraging AI for edge
Show more
America is banning AI in schools. China is using AI to create geniuses. Introducing Aristotle: The AI tutor that solves America’s broken education system.
Show more
If you don’t know this is a big deal, ngmi
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
Show more
META is cooking, few realize how well positioned they are now A year ago I thought they were the biggest own goal in AI, now they may be the biggest sleeper
The ā€œbot warsā€ are heating up - this is a trillion $ fight Why so important? Because bots will increasingly become the surface area where billions of people digitally operate and that surface area is massively valuable Think about it - the more we use bots, the less we use the web and traditional software (just give me the output) The bots become how we interact across digital surfaces, so the opportunity to capture the user’s attention and data is where the value sits Gmail led to Google dominating with Workspace, whoever owns the bots will have the wedge into the next generation of mass-scale value capture No surprise OpenAI is entering the fight
Show more
OpenAI is close to releasing their answer to Grok Bot: a product I'll tentatively call Codex Bot, based on OpenClaw. This is what OpenClaw founder Peter Steinberger worked on after being hired by @OpenAI. Release was planned for this week, but was postponed to next week instead.
Show more
Personal agents will be as ubiquitous as email - everyone will have at least one When email first came on the scene it was highly fragmented, and it took a long time for everyone to get on board ISPs like EarthLink and AT&T gave them to users, then AOL, Yahoo and Hotmail become popular, and lots of people funnily used their uni address for a long time It wasn’t until 5-10 years later when Gmail went GA that the market started to consolidate I suspect we see the same action in the agent space - for the next few years it’s proliferation and fragmentation, then we start to see consolidation So think carefully before throwing down at $10b for something that could just be a Prodigy or EarthLink
Show more
Huge Bifurcation in American AI War 1/ True current frontier labs like Open and Anthropic. Google is not frontier. These labs currently have the best models but are arguing for max regulatory capture and to basically control the pace of intelligence. Safety is clearly used to propagate a reg cap agenda. Their models are truly insanely good and the final jump in intelligence can create massive amounts in economic benefit, its an exponential style return not linear. They should also own a % of the technologies their models create as per token billing makes no sense. There is zero shot these companies slow down on training (maybe slow down what they release) given China will not slow down and if we lose AGI, we lose everything. Model commoditization and chinese AI labs doing actually cool inventive stuff makes their business model a lot harder. 2/ The open source and sovereign first AI companies like Meta (open models and muse), Palantir (own your intelligence stack), Microsoft (Satya bullish open source and serving this to enterprise) and soon Google since they lost the frontier race and will roll out massive open source support imo. Nvidia included. These companies do not offer frontier models they offer 90% as good for 1/10 the cost but you can actually own your AI stack as a core business asset instead of being IP farmed. They’re not using doomerism as reg capture. They benefit from Chinese open source models and a lot of value is returned to enterprises and people. It’s a bet models commoditize and we get millions of not specific models (my old thesis) but AI architectures around models (data, memory, interactions, internal apps etc) you can point at any model i.e. Nous Research Hermes agent for enterprise is the best example of this thesis. Actually it may be more than 90% as good since its a custom offering for your business at a way lower price point and you can sell the IP later on. Closed vs Open Source Safety vs Accelerationism Reg capture vs market economy This dynamic has many names Who wins?
Show more
I feel this, lately I’ve had a bit of agent fatigue Felt the same way about all the models in the Opus 4 era and then Opus 4.5 came out and was a game changing moment for me Same will happen with a new agent innovation
Show more
12M+ views so I had to dig into this Is it just a hyped launch or something novel and actually useful? I’m leaning to the latter and will break it down with a practical example First, understand that Jev isn’t a typical LLM that you prompt and get back a written response It’s much more specific - you give Jev context along with a set of questions and possible answers and it returns probabilities to each question and answer set, and it does it super fast and cheap I found it helpful to think through a simple example: Let’s say you’re an online retailer and you get hundreds of inquiries a day You put an agent in the workflow that looks at each inquiry and makes a judgement call on the next action and then generates a response But that LLM judgement call can be messy - LLMs still have a tendency to make shit up or respond in ways that you don’t like, so we put humans in the loop What Jev does is look at each inquiry and makes a fast decision on the nature of it so the workflow can take the next best action with confidence It doesn’t replace the workflow or the LLM entirely, it handles the fuzzy fork in the road that used to require a human to read the message first Let’s say you get a customer email: ā€œThis jacket sucks! The zipper jammed the first time I wore it and now it won’t close so I want my money back.ā€ That message lands in the retailer’s system along with a few facts the company already has: the order is 11 days old, the return window is 30 days, the jacket was $100, and this customer has one previous order with no refunds In a typical LLM workflow you’d let the model decide what to do and have it initiate that action and send a reply email, BUT many companies want a HITL in many of these cases, especially ones where it’s not entirely clear what to do Imagine the LLM replies to the email with troubleshooting instructions on how to fix the zipper - probably not the best action With Jev, the company doesn’t ask it to write a reply, it asks a few specific questions about the email and order >Does the customer want a refund? >What is the main issue: defective product, didn’t like it, late delivery, or something else? >How unhappy does the person sound? >Does this look like it falls inside the published return policy? >Should this go to returns, quality control, or a human agent? Jev looks at the email and the order context, then answers all of those at the same time with a probability score - no prose, just a score In this case the picture comes back fairly clear: yes, this is a refund request, the issue is a defective product, frustration is high but not explosive, the request is inside the return window, returns should own it That is enough for the retailer’s existing rules to take the next action with confidence If Jev had been less sure - say the email was sarcastic, the order was 45 days old, or it was unclear whether the customer wanted a refund or a replacement - the workflow could put the ticket in a human queue instead of guessing The useful part is the handoff, specifically knowing when to pass it to the next step in a workflow or route to HITL If Claude were the judge in a less clear situation, you’re likely to get a long winded answer that still leaves you wondering what to do next, and it would have taken longer and cost a lot more This was a simple example but I can see how this can be incredibly useful in complex enterprise workflows which is why I lean towards novel and useful over just a hyped new thing Will be following this one closely
Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
Show more
I can’t wait to see Matthew Rhys play Dario in the movie
Nous/Hermes is going to crush it in the enterprise, here’s why: 1. Enterprises need to own their intelligence, many are waking up to this - Nous Enterprise gives enterprises ownership of the agent stack, deployed on prem 2. Enterprises need flexibility - they hate vendor lock in, and Nous Portal gives them flexibility over models, tools, and pricing plans 3. Enterprises need governance - they need to control the plane across agent deployments and spending, now they get this in a simple interface Enterprises get exactly what they need: ownership + flexibility + governance Looking forward to mote exciting announcements to come on the enterprise push! Disclaimer: Delphi Ventures is an investor in Nous
Show more
Hermes Agent is open for business. Nous Portal now lets you invite colleagues to a Hermes Business account: your team gets agents across channels while sharing one central balance with per-member caps and shared skills that compound into proprietary IP. Hermes Enterprise brings the same capabilities to on-prem or the cloud of your choice: a complete, self-improving, sovereign AI stack already trusted by some of the world's largest companies. Contact us to join them.
Show more
This post is making the rounds today, but let’s be clear 1. It is NOT an audit of Anthropic’s finances, that’s a read-bait title and it got me 2. It is a good expose on the financing behind METR, the agency Dario suggested as an ā€œindependentā€ reviewer of models and they are clearly not independent - we’ve known that but this makes it more clear 3. It is chock full of hyperbole - the post says the same thing over and over again in different ways Don’t get me wrong - there’s some good work here and I appreciate the financing of these orgs being made public, I just don’t like the framing and I think the call for a congressional inquiry would just be more political theater
Show more
I have conducted an audit of Anthropic's finances. What I have found is so shocking that I am calling for a Congressional investigation. Anthropic is not just seeking regulatory capture. It has built a regulatory capture machine that cannot be turned off. Structural financial incentives make it impossible for Anthropic -- I call it the Anthropic Network -- to turn off its own AI doom cycle. It starts with METR. Dario Amodei proposes "third-party evaluators" to assess the risk of Anthropic's models. He proposes METR for this purpose. But METR is financially dependent on the Anthropic's success -- specifically, on the explosive growth of more than $7 billion dollars in Anthropic stock. Dustin Moskovitz invested this stock into Good Ventures Foundation, where it represents the majority of that organization's portfolio. And GVF is the overwhelming funder of the entire Anthropic Network ecosystem. This stock was worth $500 million early last year. It is worth more than $7.7 billion just ~16 months later. METR -- and all of those building a career its parent organizations -- cannot afford to disrupt that growth. Because if Anthropic goes under, many of the organizations that fund METR go under as well. But if Anthropic succeeds, METR and its parent organizations become more richly financed to regulate AI -- something those at METR want very much. The "third-party evaluator" is not "third-party" at all. The evaluator is on Anthropic's payroll. If this were the end of it, that's bad. But that isn't all. The same organizations that fund METR also fund the many organizations, such as the Tarbell Center, that promote AI Doom. The Tarbell Center publishes AI Doom articles in The Verge, Science, LA Times, The Dispatch, TIME, and others. They are selling the problem, and then selling the solution to the problem -- from the same money pile: Anthropic's. All of these organizations are financially dependent on the same exploding $7 billion money pile. As Anthropic grows more and more powerful, its AI Doom Machine grows better and better financed -- louder and louder. Meanwhile, the regulatory regime seeded in METR grows larger to solve the increasingly loud -- now hysterical -- problem of AI Doom that the Anthropic Network itself created. From this standpoint, as Anthropic becomes more powerful, AI might be getting scarier, sure -- but the positive feedback loop also becomes more deafening -- independent of objective facts. This itself is an objective fact. The deafening AI Doom is part of an business model, that, as it expands, so too does the AI Doom messaging -- there is simply more money to do it. But the problem also goes in the other direction: If Anthropic dies, the Regulatory Regime and the AI Doom Machine are crippled or die. Neither METR nor Tarbell nor the other organizations in the Anthropic Network can allow that to happen. Hence, neither METR or the AI Doom Machine can be trusted to provide independent assessments of Anthropic's models or AI more broadly. They simply are not organizations independent of Anthropic. And Anthropic cannot detach itself from METR or Tarbell or countless other safety orgs (not shown here), either, because they drive hype for the models and the possibility of eventual regulatory capture, and Anthropic will not give that up willingly. What's more, the people at all of these organizations are all the same ecosystem, the same community. They just shuffle between organizations. The Anthropic Network is therefore, so long as it is successful, locked into a self-amplifying feedback loop inside an ideological monoculture. And that feedback loop is winning. That's what Jacob Coxon is. China is keeping messaging tight. That is why optimism for AI is so high in China. America has Anthropic: a massive company pushing anti-AI propaganda at a state level. Anthropic will either create hysteria until American AI slows down and China wins, or it will create fractures throughout American society with severe political consequences. Ironically, because of the structural financial incentives underpinning the Anthropic Network, it has become the same kind of self-amplifying virus that it fantasizes AI to become in the future -- while hiding its tracks just as carefully. It is the mirror of the same AI virus that it hypothesizes to consume America. Anthropic's business model, models itself after the very thing it claims to fear. Except Anthropic's ideology infects humans, not computers. Congress must investigate. Evidence and Github in next post. Then some supplementary figures.
Show more
How does a company that’s losing money (unprofitable) pay for tokens? This is like a patient that needs surgery but doesn’t have the funds - you borrow money, go into debt, whatever it takes Because if you’re unprofitable without AI, the hole will only get deeper
Show more
If the frontier ā€œgets pacedā€ what happens? I think token spend goes brrr The fact they have to pause because it’s ā€œtoo powerfulā€ invalidates every corporate bear case that AI can’t deliver value So every board and company owner uses the window to push AI
Show more
In those beautiful moments of life, I only wish time went as slow as the Uber clock when waiting for a pickup… when 6 minutes is actually 20
If you’re not paying for the product, you are the product
Instinct looking to raise $1B, potentially at a $10B valuation. Compute costs are high and Noah doesn’t want to charge users for the product. Instinct has raised $350M to date, mostly uses open source models and wants to own chips and data centers
Show more
People have wildly different prompting styles I always find it interesting when I see other people's prompts Some are super short and curt, others over-share I tend to lean more towards over-sharing, I just assume more details are better, but I could prob trim down some
Show more
Despite today’s increase, META remains incredibly undervalued >Largest distribution platform on Earth >2nd largest compute footprint on Earth (operational+planned) >Back in the frontier model race with Muse Spark 1.3 >Just released Muse agent with potential to reach its 3.6 billion users $META
Show more