Register and share your invite link to earn from video plays and referrals.

Philip Kiely
@philipkiely
Author of Inference Engineering | Building @baseten | Not an LLM (yet)
970 Following    13.8K Followers
An open frontier is a safe frontier.
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
Show more
My agent tells me that inference engineering job openings have doubled in the past 6 months. I think that's an underestimate. Demand for inference has inflected, again. So too has demand for engineers who can operate AI models in production.
Show more
0
61
1.3K
69
Forward to community
It's the summer of Flash models. First GLM, now DeepSeek, these midsize models push the Pareto Frontier with remarkable intelligence at low prices. And, they outperform full-size last-gen models w/ new efficient architectures. Excited to see these recipes scale to 2T+ params.
Show more
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
Show more
The Blaxel team is full of wonderful people. Together, they have built a set of powerful infrastructure primitives that extend the inference layer to support agents from end to end. I'm excited to work with the entire Blaxel team in the coming months to ship incredible stuff.
Show more
Even with 1M-token context windows, a single M&A task can be 80X larger than fits in context. Great writeup about addressing this and other model limitations in a very sophisticated applied setting.
Show more
Just finished talking GLM-5.3 with Declan from Artificial Analysis. We covered: - Frontier price-performance offered by GLM-5.3 - Fighting saturation with the new AA Intelligence Index - GLM-5.3 Flash architecture and native vision I'll be around to answer questions!
Show more
Tomorrow at 11 AM PT, I'll be live on air with Declan Jackson of Artificial Analysis to discuss GLM-5.3 and GLM-5.3 Flash. We'll talk about benchmarks, performance, cost, and everything else you need to know before you switch.
Show more
An open frontier requires an open stack and open exchange. The folks working on this are some of the smartest researchers I have ever met and I am excited to see how far they can push the frontier.
Today we're announcing Base Labs, a dedicated research organization focused on advancing open-source AI. We believe in a healthy, open frontier model ecosystem. To enable this, we are working on: - Blue-sky research on continual learning, the science of RL, and how models learn across their full lifecycle, with every experiment and recipe shared openly. - The BaseHub Data Foundry: the highest-quality open RL environments, training data, and real-world benchmarks, built for anyone to train and benchmark on. - Post-post training: taking open-source models and making them better, safer, and more aligned through continual post-training, built on our research, and deployed with our frontier safety stack so organizations can use open-source models with confidence. - Making models cheaper and more performant through our model performance research. This is a mission-driven research effort, not a commercial product. We believe the health of the open-source AI ecosystem matters and that the best way to advance it is to do science in the open. We’re hiring engineers, researchers, and research fellows to advance this mission.
Show more
Imagine how cool it would be to watch NFL with this tech -- first person view as the quarterback or any of the 22 guys on the field.
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
Show more
Hanging out next Thursday 9/10 with my friends from NVIDIA Dynamo and SGLang. We'll have tech talks on inference optimization at every layer of the stack. Join us!
Teach your agents inference. The full text of Inference Engineering is now available to LLMs everywhere. No more parsing PDFs. 3,000 tokens | links to each chapter: 70,000 tokens | full text:
Show more
Faster kernel != faster model. Great breakdown of systems thinking around kernel optimization.
I may have taken the wrong lesson from Bryan Johnson's visit to our office. If Bryan has achieved such incredible results after getting serious about health in his 40s -- and if the tech is only getting better -- maybe I can wait another 10 or 15 years before investing in my own health. This pattern of thought comes up often in the AI industry. If models are getting better, why waste time making incremental progress on the infrastructure and application layers when growth in model capabilities will eventually consume all progress? After nearly five years working in AI, I am confident that there is a lot of value to deliver along the way between improvements in model quality. And that continuously operating at the frontier positions you to take proper advantage of these advances. Bryan makes the same argument for health. Small habits and choices made early compound for a lifetime, and putting the essential structures and routines in place makes the cool optimizations actually worthwhile. Thank you to the entire Immortals team for a fantastic event.
Show more
Rapid changes in capabilities require rapid changes in abstraction. Alex & co offer a compelling vision for the next layer of AI tooling.
A small team at Baseten, led by the official @waterloo_intern, made Wan 2.2 video generation 53.6x faster. Often, optimization is about making existing things better. This is a case of performance work unlocking a completely new class of capabilities across real-time video.
Show more
Inkling by Thinking Machines is a new open model with 975B parameters and omni-modal inputs. I tested it across coding, tool use, and vision by using it to build and operate Ink, a digital calligraphy coach.
Show more
Every company worldwide must own their intelligence.
Most companies today are renting their intelligence a token at a time. That's @baseten @philipkiely's read, and it's the take from our conversation I wanted to highlight. AI equals @OpenAI, @AnthropicAI, @GeminiApp tokens flowing into whatever you're building. Which means the thing you're built on isn't yours. A price hike, a deprecated model, a rate limit you didn't see coming and your product breaks on someone else's schedule. You don't own your roadmap. You're leasing it. The companies Philip watches are done with that. Fine-tuning their own models, running open weights, owning their tokens end to end. The GPT wrapper era is over. He's blunt about it: cool in 2023, uncool by 2024, dead now. What replaces it is companies that own the thing they're built on. His other take:"I am not worried about an AI bubble." I didn't have to take his word for it. Since we filmed, the receipts came in. @Baseten raised a $1.5B Series F at a $13B valuation — their fourth raise in 18 months, inference volume up 40x in a year. The line in their announcement: Their customers are "more motivated than ever to own their intelligence." Same phrase Philip used with me on the pod. The companies that own their intelligence just got $1.5B of validation. The ones renting it got a deadline. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Intro (02:30) What Inference Actually Is (05:00) Why You're Renting Your Intelligence (08:00) OpenEvidence & When Latency Means Lives (11:00) Writer Who Engineers, or Engineer Who Writes? (15:00) Joining Baseten Before ChatGPT Shipped (19:00) What He Got Wrong About Inference in 2022 (23:00) GPT Wrappers Are Dead. Good. (27:00) Writing 256 Pages in a Field That Moves Weekly (32:00) The Chapter to Read With the Most Skepticism (36:00) The 24 Hours After DeepSeek Drops (42:00) Why 8 GPUs Is the Magic Number (47:00) Rent vs. Own: Where the Break-Even Sits (51:00) "I Am Not Worried About an AI Bubble" (56:00) Every Wafer Gets a Chip Pressed Into It (01:00:00) Will NVIDIA's Moat Ever Crack? (01:05:00) Is There a Black Market for GPUs? (01:09:00) Inference Engineering as a Career (01:13:00) There Are Only Two Kinds of Models This is a @Composio "Agents at Work" podcast, where I chat with founders building the next leap of AI. Follow for more :)
Show more
Technical writing is the #1# top-of-funnel motion at a $12B company @philipkiely broke down exactly how he does it, his full writing process, the hindsight 20/20 lessons, and how one article drove 500K+ views This [technical] Write and Learn workshop #5# is part of our Independent Studies series, hosted with @swyx @KernelLabs_ai. All sessions are recorded. Our next one will be in two weeks :)
Show more
Local and datacenter inference engineers face different constraints, but we have a lot to learn from one another. Thanks for the conversation Sero!
Here's my conversation with Philip Kiely. A deep dive into inference engineering, his book his available online for free, and it's worth reading. Inference is becoming more important to people and organisations by the day, learning about the topic will prepare you well.
Show more