Register and share your invite link to earn from video plays and referrals.

Baseten
@baseten
Inference is everything.
79 Following    16.4K Followers
Kimi K3 is now live on our Model APIs, day 0.
Thrilled to double down on our partnership with @baseten @tuhinone is one of the clearest thinkers I know on AI infrastructure. He sees a cross section of how AI is actually being used across the leading AI natives, and his job is making them fast, reliable, and cheap at scale. We discuss why inference is becoming one of the biggest markets in software, what happens as open-source models converge with frontier, and a peek into future AI workloads! Enjoy my conversation with Tuhin:
Show more
New open-weight models from the US! We're proud to continue partnering with @eisokant and the @poolsideai team on inference.
Another big day for American open-source from the team at @poolsideai! We're proud to be your partner for inference.
The American open-weight ecosystem is growing. We're thrilled to support Poolside as they push open-weight coding models forward. Poolside's flagship model family is Laguna: open-weight agentic coding models built for long-horizon tasks. Today, they're releasing Laguna S 2.1, a 118B (8B active) MoE model with up to 1M context. @Madisonkanna sat down with @poolsideai Co-CEO @eisokant to talk about the launch, and what it means for the ecosystem:
Show more
Big day for American open-source AI. For the launch of Laguna S, I sat down with @eisokant to discuss its architecture, the economics of open weights, and the question of who gets to build intelligence. Timestamps: 0:00 Intro 1:50 Why Poolside started opening its models: the oligopoly on intelligence 4:28 Getting nerd-sniped by Karpathy, building LLMs before anyone cared 11:20 Laguna S: 118B parameters, 8B active, built in 8 weeks 13:05 The future of software engineering: behaviors over IQ 14:25 Sliding window attention, 1M context, the model factory 15:50 Being an American open-source lab 20:47 The economics of open weights 24:22 Who gets to build intelligence? The 12–18 month window 30:28 Erdős 397 in 30 minutes 31:48 Extracting transcripts with a debugger Congrats to the Poolside team on the launch!
Show more
You can now generate 5 seconds of video in under 2.5 seconds, courtesy of the Baseten kernels team.
A small team at Baseten, led by the official @waterloo_intern, made Wan 2.2 video generation 53.6x faster. Often, optimization is about making existing things better. This is a case of performance work unlocking a completely new class of capabilities across real-time video.
Show more
"So, can a model learn facts continually in its weights? Creating usable knowledge: solvable Preserving capability: solvable Keeping earlier facts reachable: unsolved" New research from our Head of Training just dropped on arXiv.
Show more
1/ Can you actually get new facts into an LLM's weights without breaking the model? This question decides how we approach continual learning: should memory live in the context (retrieval, compressed caches) or in the weights themselves? We spent a long time measuring it, and it breaks somewhere much stranger than we expected, making us much more bullish on compressed kv caches and ICL for continual learning, as opposed to weight updates themselves 🧵
Show more
We have dozens of open roles across Toronto, Montreal, New York, and San Francisco. See the full list here:
We're building in Canada! We opened two new offices in Toronto and Montreal, and we're growing fast. If you’re excited about building high-performance AI infrastructure for the world's most ambitious companies, we'd love to meet you!
Show more
I bet against @part_harry_ when he said he could add vision to GLM 5.2 with projector only training. Never bet against the quants, training a 2 layer MLP is all it took. We're going to open source this! Many more of the team's hobby horse projects coming out soon.
Show more
GLM 5.2 is one of the best open models available, but it can't support image inputs natively. Until now.
the most underrated line in the inkling blog: "[it's] not the strongest overall model available today, open or closed." thinky ships a 1T MoE with 1M context and native multimodal without benchmarkmaxxing the entire bet is that post-training scales. great american open models are so back. a good day for anyone in the app layer looking to build a moat with RL
Show more
Inkling by Thinking Machines is a new open model with 975B parameters and omni-modal inputs. I tested it across coding, tool use, and vision by using it to build and operate Ink, a digital calligraphy coach.
Show more
Big day for American models! Excited to partner with @thinkymachines to provide day 0 support for Inkling.
Inkling is live on our Model APIs! Thrilled to be a day 0 launch partner for @thinkymachines.
The web's next biggest users will be agents, and @p0 is building for that future. We're proud to be their partner for training and inference!
Some workloads demand either the highest throughput or the lowest latency. Embedding workloads need both. We built Baseten Embeddings Inference (BEI) to meet that need, and we're thrilled to partner with @turbopuffer to power BEI-optimized models in tpuf!
Show more
now in beta: native embeddings in tpuf embedding is the most painful part of puffing. we want to make it easy you can now convert chunks to vectors as you read and write to turbopuffer, without extra calls to an embedding model provider API docs:
Show more
Really enjoyed this. We covered why I lasted only a few days into an Oxford PhD, why you should learn RL by touching nothing but the config and watching the curve go up, and why the intelligence ceiling of a specialised open-source model now subsumes the frontier for most real tasks. Also the story of negotiating with @tuhinone in my pajamas
Show more
Research has always been core to Baseten, from training to model performance and beyond. Charlie, who leads part of our training team, sat down with Madison to talk about what it takes to be an AI researcher, from his PhD at Oxford to founding a company dedicated to post-training (and joining Baseten!)
Show more
How to become an AI researcher with @oneill_c Charlie co-founded Parsed to build specialized open-source models that can outperform frontier labs. I first met Charlie when Parsed was acquired by Baseten, and now he leads our model development team. Charlie is one of the smartest people I know, and I had the pleasure of talking to him about: 0:00 Intro 3:13 Leaving Oxford to start a company 6:37 Becoming an AI researcher 15:37 Developing a unique POV as your moat 22:04 Parsed origin story 26:01 Big Token, the case for open-source models 33:40 Post-training, fine-tuning, specialization 46:52 Will open models catch up with closed models? 51:50 AI-led job replacement vs job creation 54:45 How to get into inference engineering This is one of my favorite conversations I’ve had in a long time. Made with @ad0rnai behind the scenes. Enjoy!
Show more
Try Step 3.7 Flash by @StepFun_ai: Learn more on our blog: