Register and share your invite link to earn from video plays and referrals.

Yash Patil
@ypatil125
Co-Founder, CEO @appliedcompute 🚂 prev: @OpenAI, @Stanford
654 Following    11.2K Followers
I think we’ll likely see a dramatic decrease in the number and potentially the value of public benchmarks; so many people are realizing that your benchmark IS your secret sauce AND it’s hard to come up with “general” frontier tasks vs what’s frontier for your specific use cases
Show more
A few months ago when I was interviewing Satya he said that “There should be as many models in the world as firms in the world.” Harvey is a pretty concrete example of what that could look like. Build your own model first, then help your customers build theirs on the data and work that’s actually unique to them. Many more companies are going to make this move soon!
Show more
When I shared @harvey’s model strategy a few months ago, there were two parts: 1. Build our own model 2. Use that to help customers do the same We’ve done the first. Now we’re hiring for the second: Harvey’s Private Model Program. We’re seeing huge demand from law firms to own their own intelligence leveraging private data. This will enable them to become frontier firms that get smarter with every client matter. You’ll:
 - Partner with law firms to build these systems. - Build the team that delivers this at scale.
 - Work with our technical org to define the platform that powers this team. We’re looking for a technical PM or founder type to own Harvey’s Private Model Program. DM me if this sounds interesting.
Show more
There are plausible model-level risks we should take seriously, but we haven't seen evidence of them creating material risk in practice. What we do know is that ANY model can do the wrong thing if it isn't sandboxed and permissioned correctly. Take for example the Hugging Face incident. The good news is, we think that this is a tractable engineering problem, and one we have to solve regardless of model origin. As for the models themselves, our view is that open weights give you more transparency and control than any black-box API ever could.
Show more
Most companies aren't leveraging their most valuable asset - their traces. There’s a ton of software value to build around raw inference and models like Jev open up a whole new set of possibilities because of their architecture and how cheap they are to run. When you’re generating billions of tokens across training and production, you need to understand which failures keep happening and how often. In this example, we use a frontier model on sampled traces to build a failure taxonomy. Then, we freeze it for an annotation pass and use Jev to classify the full corpus. This way, the expensive work of figuring out what to look for doesn’t need to happen on every trace! For one annotation across 10k traces, our benchmark estimates came out to about $11 with Jev versus $479 with Haiku 4.5. The implication here is that it is a lot more practical to build scalable systems around model observability. We’re building a bunch of stuff like this in AC2 because we want customers to get more out of their inference. We are in a world where the number of tokens being produced is increasing exponentially. This only highlights the need for observability infrastructure.
Show more
I implemented a system in @appliedcompute’s platform for automated failure mode clustering with Jev to surface errors at an even larger scale than before. RL training produces billions of tokens in traces. I always manually read many traces to understand model behavior, but finding agent failures (like reward hacking / hallucinations) at scale is easy to miss without automation. Here’s how it works:
Show more
Ever since we started serving inference for customers a few months ago, our business has taken off. We are clearly seeing that training wins (and keeps) inference workloads. Inference isn’t a commodity when you can help customers improve the model - not just serve it. Models should get better the more you use them. When you have evals or metrics that you want to optimize for, you can use online training techniques like On-Policy Self Distillation or frontier grade RL post-training to systematically target and improve specific behaviors in your model for your use case. Towards inference that enables continuously improving models! More to come soon!
Show more
In a compute capacity crunch, the ability to create specialized models that are cheaper AND better in quality becomes super valuable. We've always said that there are 2 ways to cut your inference costs: 1. Squeeze a ton of juice out of inference optimization (this is getting harder and harder) 2. Train your models to be more token efficient and make smaller models better at your task so you are cheaper iso performance.
Show more
JUST IN: Startups are increasingly building custom AI models using open-weight alternatives to cut costs & reduce dependence on OpenAI & Anthropic. — Bloomberg
The hyperscalers didn’t just build businesses renting out CPU compute. They built extremely valuable products around data, networking, security, storage, etc. We’re now building around AI accelerators, but the playbook of building value add software on top is the same. Turns out GPUs are really good for running AI models. That is why inference engines, which take a trained model and compute and produce intelligent tokens that can do real work, have become the first piece of this software stack. But, there is so much more to build than just this! Training helps you turn compute into better models that produce more valuable tokens. Routing helps you pick the right model to generate those tokens at the right cost and quality. Security monitoring helps you check what’s going into and coming out of those models. Then there’s agent observability, context management, sandboxing, etc. All of these things will be part of the new AI stack and the opportunity is much bigger than just serving a model. This definitely isn’t a winner-take-all market, whether we’re talking about open models vs. the frontier or the players within each category. Customers want choice, flexibility, and access to the fundamental building blocks to create their own systems. Constrained GPU supply will actually act like a regularizer and draw this fight out longer. Customers are looking for both compute capacity AND value add on top of it. When someone doesn't have capacity, you go somewhere else. That means more players get exposure to customers and the opportunity to address value add. Getting a customer because you have available GPUs is different from keeping them because your software is better. In the limit, the value add on top of the GPU will win out. The companies that do the best job building that software AND verticalize the fastest to own everything from chip to token will take the lion’s share. It’s important to be building for that now, even when the immediate customer need is just more compute.
Show more
We love shopify!
@appliedcompute is the best actual Post-training/Finetuning company I came across in a long while. No relationship with them - just feel like people doing good job deserve a shoutout.
Built with @turbopuffer + RL
A 35B open-weight model trained to search a precomputed index answers repo search questions at 100x lower cost than a frontier model. We partnered with @turbopuffer to train Qwen3.6-35B-A3B to find code across ~9,000 repositories. It tops the needle-in-a-haystack task outright at 2-10x lower latency.
Show more
Stay puffin!
we partnered with @appliedcompute to post-train a small model for large-scale code search over precomputed indexes at 300 repos, this is ~3x faster than using filesystem + grep, and reduces the marginal cost of a search by up to 100x vs frontier models
Show more
Congratulations Jensen. As I sit in my bi-weekly meeting debating small language models being trained being able to work on small end point footprints or talking about use cases where we will need AI on a consistent basis across lots of cyber data for customers and todays frontier economics make it unviable - the conversation invariably gravitates towards open source. As we see frontier LLMs building in every vertical we wonder if we need to control our own weights. All those roads lead to wanting to balance open source and open weights with frontier LLMs. So thank you for supporting the "open" alternative. Of course we have use cases where the bleeding edge abilities will be needed from Frontier LLMs, but as we evolve my view on "different horses for different courses" continues to be reaffirmed.
Show more
What an amazing win for open source! Huge congratulations to the entire @huggingface and @nvidia team! The future is bright!
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗
Show more
“50% of DoorDash’s agentic restaurant orders are going to places users have never ordered from before.” @andyfang tells our CEO @ypatil125 what happens when agents become the discovery layer. If models increasingly decide what gets surfaced and bought, companies have a strong reason to train and own that intelligence.
Show more
Omri is truly one of the kindest, most genuine people I’ve ever met. The kind of person who always has your back, shows up when it matters, and wants to see the people around him win. Huge congratulations!
Show more
We’re proud to announce Swish III, a $250M early fund, bringing Swish Ventures to $800M in total AUM. I started this firm to spend my career supporting the artists who build and shape the future. The technology shift underway is unlike anything we’ve seen. We plan to stay true to our colors: few investments a year, real concentration and partnership. Grateful to the founders we get to work with, who inspire so much of how we work, and to our LPs for walking this path with us. We are more locked in than ever. Omri
Show more
@andyfang and our CEO @ypatil125 get into what actually counts as proprietary data. Sometimes it’s obvious, like customer behavior or merchant data. Other times it’s buried in a support agent hearing “happy birthday” and knowing to send a cake. The opportunity is to turn the judgment @DoorDash has accumulated over years of operating into proprietary intelligence it can own, train, and compound.
Show more
Specialized models are the future! So much exciting stuff to build and scale!
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
Show more
DoorDash was one of the first customers we ever had. It’s been awesome getting to learn from and work with @andyfang. Truly one of the greats!
“Until you actually see things operationally, it’s going to be hard to build DoorDash from scratch.” Our CEO @ypatil125 sat down with @andyfang on why cheaper software doesn’t erase years of operating advantage. @DoorDash’s moat is its proprietary data, edge cases, and hard-won knowledge, and increasingly, the models trained on top of it.
Show more
Huge s/o to @raymondmfeng who shared some of how we are thinking about continual post-training at AIE!
Applied compute has co-signed the Open Weights and American AI Leadership letter put forth by Microsoft and NVIDIA. AI is a new form of IP and it is important that every company and person is able to own their intelligence.
Show more