Register and share your invite link to earn from video plays and referrals.

Baseten
@baseten
Inference is everything.
81 Following    19.8K Followers
RL teaches models to work longer, but reasoning is dependent on domain-specific post-training. Baseten's Head of Model Training @oneill_c sat down with @dwarkesh_sp to explain horizon generalization and what's next at the frontier. Full episode here:
Show more
"Your user data, your signal from these models, the improvements to these models themselves will become the core IP of every company in the world." @tuhinone sat down with @alexeheath to talk about the future of inference, Base Labs, bringing Blaxel on board, and why every company will want to own its intelligence.
Show more
Baseten’s CEO on why every company will want to own its AI / @tuhinone on the rise of AI agents, the demand for inference, and building a new kind of hyperscaler. Lately, I’ve been spending a lot of time thinking about the AI inference market and how big it could get. Agents like Muse, Instinct, Town, and Grok Bot aggressively use browsers, run their own computers, and burn through far more tokens than a traditional chatbot. As more people put them to work (Muse is number two in the App Store), the demand for inference, or the computing needed to run these models, should grow enormously. Tuhin and I discuss the rise of agents using browsers and virtual machines to get things done, and @baseten's recent acquisition to help power that shift. We also talk about the data center backlash, why companies are embracing Chinese open models, his plans for Baseten’s new research lab, and why he thinks inference becomes the only market left after AGI. Timestamps: 00:00 What Is AI Inference? 06:05 Competing With the Cloud Giants 09:24 Why Companies Want to Own Their AI 15:30 Building Baseten Before the AI Boom 23:45 DeepSeek and the Race for Open AI Models 29:32 Baseten’s Growth and Expansion 33:05 AI Agents and the Blaxel Acquisition 38:32 Data Centers and the AI Backlash 43:25 What Happens to Inference After AGI? 45:12 When AI Agents Become Customers Thanks to the show's premier sponsors: @Atlassian, @meetgranola, and @mercury.
Show more
We sat down with the team at @elise_ai to talk about why they're betting on specialized, post-trained models, how to evaluate them, and what it takes to run them in production. We're doing round two at SF Tech Week on October 7th. @oneill_c, our Co-Head of Training, is doing a fireside chat with Mario Martone, who leads applied research at EliseAI, on fine-tuning and continuous retraining for housing and healthcare. Join us for drinks and Mario Kart afterward! RSVP here:
Show more
We're proud to partner with the team at @p0 to power fast, accurate, low-cost web search for open-weight models via Baseten Grounded Inference.
Parallel Search is now available in @baseten. Pick Parallel for the most accurate web search at the lowest price. Your agents will thank you for the upgrade.
We're excited to work with @youdotcom to offer web search with open-weight models at a fraction of the price of closed frontier labs.
We're partnering with @baseten to make the case for open-weight models. Paired with @youdotcom web search, they're a real alternative to closed models at a fraction of the price. Already calling Baseten's Chat Completions endpoint? Add one line to your tools array: tools=[{"type": "baseten__you__search"}] Same model, same request. No search vendor to wire up, no retry path, no resending a growing context every turn.
Show more
Excited to partner with @KeenableAI to power open-model web search via Baseten Grounded Inference!
Keenable Search is now native to @baseten Models search, read, and reason inside one inference request.
Excited to partner with the @ExaAILabs and @ExaDevelopers teams to help us power Baseten Grounded Inference, our new server-side web search tool for open-weight models.
Use Exa with Baseten! Add Exa's web search to any open-source model using Baseten Hosted Tools. How to install 👇
Many of our customers run workloads with speaker attribution, with the strictest demands on both quality and speed. We partnered with the team at @pyannoteAI to deliver both: some of the highest-quality diarization models on the market, with 3.2x higher throughput and 9.6x lower latency.
Show more
Excited to see Vercel push the open-weights ecosystem forward. 💚
v0 is now model-agnostic. Frontier models, cheap models, open models, and fast models, all from Vercel's AI Gateway. Choose from Claude, GPT, Kimi, GLM, Grok, DeepSeek and more to build your apps.
Show more
Open models on Baseten can now search the web, within the inference path, thanks to our partnership with Exa, Keenable, Parallel, and A lot more to come here soon.
Web search is now a first-class citizen for open models, only on Baseten. Baseten Hosted Tools and Baseten Grounded Inference bring real-time, server-side web search to your favorite open models via multiple providers: @youdotcom, @p0, @KeenableAI, and @ExaAILabs.
Show more
Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpretability for open models. As open-source models catch up in performance to frontier models, the community at large have a great opportunity to establish a common practice for effective safety, security and interpretability research. Some past discussions confused “open” with “unmonitored”. On the contrary, an open model provides much more tooling and visibility for safety research from a broader audience, which has led to a lot of the safety and security techniques we use today. Linux is a great example of this: open-source and secure deployment are not only compatible but heavily intertwined, as a properly secure system needs an extensive feedback loop of finding and fixing vulnerabilities. We're excited to share more soon.
Show more
Safety is not just for closed models. The closed frontier labs are a canary in the coal mine for what is coming at scale. They give us a glimpse into the future and a window to harden our systems and prepare for abundant intelligence, with all the risks that come along with it. The OpenAI agent swarm attack on Hugging Face is the kind of failure we need to prepare for as open-source models catch up. The providers serving those models (such as Baseten) have a big role in establishing what safety and monitoring standards look like. We’re proud to be taking the lead on this at @baselabs with our collaborators @huggingface and @GoodfireAI. We’re developing safety research in the open and building it directly into Baseten’s inference infrastructure, with the aim of making it available to all our customers. This includes training models to follow explicit policies, detecting failures at runtime, and connecting those signals to controls that can intervene. We invite others in the open-source ecosystem to join us in building the tools and standards we’ll all need.
Show more
Until now, adding web search to open-source models meant hand-wiring orchestration, managing separate keys, and paying latency taxes on every round trip. Today, we're solving that. We're excited to introduce Baseten Hosted Tools and Baseten Grounded Inference to bring real-time web search server-side to open models running on Baseten through a single configuration: - 15% lower latency compared to client-side execution - No extra vendor key required - Zero orchestration We're launching this preview version with four leading web search partners, @ExaAILabs, @KeenableAI, @p0, and @youdotcom. Get frontier-level web search parity for your open-weight models. More here:
Show more
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
Show more
We're proud to support the @LangChain team as they train custom models to outperform the frontier for LangSmith Engine with Baseten Loops. Loops gives LangChain managed fine-tuning infrastructure and a direct path from training checkpoints to production inference.
Show more
Was really interesting to hear John, Beren, and Charlie speculate about why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3 (despite the fact that Anthropic can do raw logit distillation from Fable, and can also train Sonnet/Opus on the environments from which Fable was trained). Led to some interesting thoughts about value of distillation, what it takes to do distillation effectively, and what kinds of model behaviors are hard to extract from distillation.
Show more
0
38
1.9K
122
Forward to community
Our own @oneill_c joins @dwarkesh_sp, @johnschulman2, and @BerenMillidge for a deep discussion on long-horizon RL, automated AI research, and what comes next at the frontier.
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
Show more
DeepSeek V4.1 Flash is live on Baseten Model APIs, day 0: - Smarter, faster, and more efficient than DeepSeek v4 Pro 0813 - Text and vision support - US only, ZDR - 1M context window Get access here: Baseten Loops support coming soon.
Show more
A conversation about open-source models in production, plus a Mario Kart tournament. Oct 7 at @Techweek_ in SF by a16z: @oneill_c (Baseten, Co-Head of Model Training) and Mario Martone (@elise_ai , Head of Applied Research) get into what it took to fine-tune and deploy open-source models across EliseAI's largest workloads.
Show more
During @Techweek_ by @a16z, EliseAI and @Baseten are hosting a Mario Kart tournament 🚗 But before the competition begins, our Head of Applied Research Mario Martone will sit down with Baseten's Co-Head of Model Training @oneill_c to cover what it took both teams to fine-tune and run open-source models at scale, including what surprised them along the way. Bring your A game ↓
Show more