Register and share your invite link to earn from video plays and referrals.

Philip Kiely
@philipkiely
Author of Inference Engineering | Early @baseten | Not an LLM (yet)
888 Following    10.3K Followers
A small team at Baseten, led by the official @waterloo_intern, made Wan 2.2 video generation 53.6x faster. Often, optimization is about making existing things better. This is a case of performance work unlocking a completely new class of capabilities across real-time video.
Show more
Inkling by Thinking Machines is a new open model with 975B parameters and omni-modal inputs. I tested it across coding, tool use, and vision by using it to build and operate Ink, a digital calligraphy coach.
Show more
Every company worldwide must own their intelligence.
Most companies today are renting their intelligence a token at a time. That's @baseten @philipkiely's read, and it's the take from our conversation I wanted to highlight. AI equals @OpenAI, @AnthropicAI, @GeminiApp tokens flowing into whatever you're building. Which means the thing you're built on isn't yours. A price hike, a deprecated model, a rate limit you didn't see coming and your product breaks on someone else's schedule. You don't own your roadmap. You're leasing it. The companies Philip watches are done with that. Fine-tuning their own models, running open weights, owning their tokens end to end. The GPT wrapper era is over. He's blunt about it: cool in 2023, uncool by 2024, dead now. What replaces it is companies that own the thing they're built on. His other take:"I am not worried about an AI bubble." I didn't have to take his word for it. Since we filmed, the receipts came in. @Baseten raised a $1.5B Series F at a $13B valuation — their fourth raise in 18 months, inference volume up 40x in a year. The line in their announcement: Their customers are "more motivated than ever to own their intelligence." Same phrase Philip used with me on the pod. The companies that own their intelligence just got $1.5B of validation. The ones renting it got a deadline. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Intro (02:30) What Inference Actually Is (05:00) Why You're Renting Your Intelligence (08:00) OpenEvidence & When Latency Means Lives (11:00) Writer Who Engineers, or Engineer Who Writes? (15:00) Joining Baseten Before ChatGPT Shipped (19:00) What He Got Wrong About Inference in 2022 (23:00) GPT Wrappers Are Dead. Good. (27:00) Writing 256 Pages in a Field That Moves Weekly (32:00) The Chapter to Read With the Most Skepticism (36:00) The 24 Hours After DeepSeek Drops (42:00) Why 8 GPUs Is the Magic Number (47:00) Rent vs. Own: Where the Break-Even Sits (51:00) "I Am Not Worried About an AI Bubble" (56:00) Every Wafer Gets a Chip Pressed Into It (01:00:00) Will NVIDIA's Moat Ever Crack? (01:05:00) Is There a Black Market for GPUs? (01:09:00) Inference Engineering as a Career (01:13:00) There Are Only Two Kinds of Models This is a @Composio "Agents at Work" podcast, where I chat with founders building the next leap of AI. Follow for more :)
Show more
Technical writing is the #1# top-of-funnel motion at a $12B company @philipkiely broke down exactly how he does it, his full writing process, the hindsight 20/20 lessons, and how one article drove 500K+ views This [technical] Write and Learn workshop #5# is part of our Independent Studies series, hosted with @swyx @KernelLabs_ai. All sessions are recorded. Our next one will be in two weeks :)
Show more
Local and datacenter inference engineers face different constraints, but we have a lot to learn from one another. Thanks for the conversation Sero!
Here's my conversation with Philip Kiely. A deep dive into inference engineering, his book his available online for free, and it's worth reading. Inference is becoming more important to people and organisations by the day, learning about the topic will prepare you well.
Show more
Its Inference Day at @aiDotEngineer World’s Fair! Catch my talk at 1:30 on the Inference Track (room 2016). I’ll cover what has happened in inference since I published the book. And grab your copy of Inference Engineering at the Baseten booth.
Show more
Down a man, but never a doubt USA 2-0
The conviction to adopt open models like GLM-5.2 comes from building clear and specific evals. This is a great example of what that looks like in practice.
@baseten Read the full GLM-5.2 eval →
Inference Engineering is flying off the shelves at @aiDotEngineer world’s fair! Stop by the Baseten booth for your free copy.
Deeply technical kernel work by Ali. Really excited about fast video inference.
Paper copies of Inference Engineering are temporarily sold out online, should be restocked early next week. Will have some final copies from this print run with me at AI Engineer World's Fair in SF!
Show more
This is the perfect intersection of two current trends in inference: 1. Live, dynamic updates of inference systems instead of static configuration 2. Using training to improve inference performance, not just raw intelligence
Show more
People really like GLM-5.2
GLM 5.2 is already top 10 on OpenRouter this week, sitting right next to Opus 4.8. it's also been really strong in Deep Agents across long-horizon coding, big context, tool use, and verifying work before it says it's done. open model and much cheaper to run than its peers. On par with GPT 5.5 speed if used with @baseten (>280 TPS and <0.8s TTFT) super easy to try for yourself:
Show more
I was asked a great question: "what companies are doing a great job with community these days?" I don't have a good answer. Help?
Mostly survived the Corey Quinn examination!
In this thread I will set up @baseten since they raised an absurd amount of money today and probably should spend some of it on Shitpost Crisis Response. Plus @mikejulian sent me an invite:
Show more
We are at the beginning of the most extraordinary market opportunity in history. The world's most sophisticated AI builders have realized that owning intelligence with domain-specific post-training and dedicated inference infrastructure is the path to building durable, differentiated businesses. Soon, this thesis will be obvious everywhere. Working with exceptional customers like Abridge, Clay, OpenEvidence, Harvey, and Notion, we've built the defining system for refining user signal into custom intelligence and delivering it at scale. I love working at Baseten. If you are amazing at what you do, you will love working here to. We are hiring across engineering, research, and GTM. DMs open.
Show more
1M context at these speeds 👀
@baseten model performance team is absolutely cracked. @Zai_org GLM 5.2 is now 4x faster running at full 1M context! Already available to use in your favorite coding harnesses, here it is COOKING in @FactoryAI Droid and @opencode Docs for how to get it in comments
Show more
Great inference requires a great model Great models require great data Great data requires capturing what actually happens in production Enjoyed chatting with the @JudgmentLabs team about everything from agents to GTM strategies (ice cream is surprisingly high ROI)
Show more