Register and share your invite link to earn from video plays and referrals.

Vipul Ved Prakash
@vipulved
Co-founder, CEO @togethercompute
1.1K Following    8.5K Followers
Our serverless APIs for open weights models have become exquisite machinery (that we plan to write more about soon!) One effect is that @togethercompute now one of the largest originators of open tokens in the US, with a wild week over week growth curve that reflects the pace of OSS adoption. Will serve 1.8T tokens today on @OpenRouter alone. You can get started with $5 —
Show more
Nice stats from @OpenRouter that illustrate how @togethercompute delivers solid and scaled performance for agentic workloads. We are serving GLM 5.3 and GLM 5.3 Flash @ top decile of TPS, latency, cache rate, and doing it at large volumes, 23% and 30% of all OpenRouter traffic, and OpenRouter is a fraction of our overall API traffic. Running scaled inference for agentic applications is a lot more than maxxing a single metric ... you need to optimize on all dimensions and doing so at scale with reliability.
Show more
Getting prepped for producing tokens on Vera Rubins! We introduced VR support in ThunderKittens today.
A primer from @togethercompute on how to bring open source models to work alongside or substitute closed models.
My first byline for @CNBC, co-reported with @KatieTarasov - how are legacy data center operators like @Equinix fitting into the AI boom? It's all in the inference (and a partnership with @nvidia and @togethercompute).
Show more
At Equinix Horizon, Adaire Fox-Martin asked Jensen Huang why inference needs to sit close to your data, not just more compute. Today that became Equinix® Inference Exchange, built with @nvidia + @togethercompute.
Show more
AI inference has a location problem: where should workloads run, and how should they connect? Equinix® Inference Exchange brings together #Equinix#, @nvidia and @togethercompute to help enterprises move from AI experimentation into production. Learn more:
Show more
We are partnering with @Equinix and @nvidia to build a global inference fabric for inference. This partnership brings inference (literally) close to enterprise data and applications through our new inference edges colocated in Equinix's distributed, enterprise-class, low-latency datacenters.
Show more
Open source is becoming the default way enterprises build with AI. Not just which model runs, but where it runs. Today we're partnering with @Equinix and @nvidia on Equinix Inference Exchange: our open-model platform, live across Equinix's global data centers. Model choice and performance were never a trade-off.
Show more
We executed a landmark partnership with @HUMAIN today to offer open source models on 250 mw (-100K+ chips) of capacity over the next 12 months. The demand for open and custom AI tokens is vertical and this partnership represents one of the largest so far focused on delivery OSS tokens globally.
Show more
We just signed one of the largest AI infrastructure deals for open source, period. 250MW data center. $5B+ in annualized revenue. Built with @HUMAIN in Saudi Arabia.
Excited to share that 0x Alpha will be on @togethercompute as early as tomorrow morning!
Excited to partner with @togethercompute as we scale Higgsfield’s inference infrastructure
Alex (@alexmashrabov) has built a rocket ship @higgsfield_ai. Excited to be partnering with them @togethercompute to power their inference backend!
Huge welcome today: @higgsfield_ai, the leading AI video and image creation platform behind Cinema Studio, is now a Together AI customer, fresh off their $400M Series B. Cinematic intelligence, one unified workflow, studio-quality video at any scale. Their video models run on Dedicated Container Inference, built for long-running, multi-GPU jobs: autoscaling, queues, traffic isolation, retries, monitoring. Proud to power the inference behind it.
Show more
Cool to see @togethercompute top the list of fastest growing vendors of summer 2026!
Our summer fastest-growing vendors list just dropped, and the story is: infrastructure won. Turns out the real AI hype isn't the app, it's what's under it.
We recently introduced support in @togethercompute endpoints for seamless A/B experiments and rollouts. Our customers are using this to rollout new weights safely and rapidly for small to large deployments.
Show more
Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. A/B testing belongs at the endpoint, not in your app code. Same endpoint name, API, and keys for your clients. No feature flags, no hash-mod-100 in client code, no spreadsheet explaining what group A vs B means. Split a live endpoint's traffic into one control and up to 20 variants, each with a fixed percentage. Ramp with a single call. Delete the experiment and 100% of traffic returns to the control, with nothing left to unwind. Read the full walkthrough:
Show more
DeepSeek V4 Pro 0813 is live on Together AI DeepSeek’s flagship V4 Pro release brings a 1.6T MoE architecture, 1M context, and three reasoning modes for coding, agents, and complex reasoning.
Fantastic initiative by @nvidia. “Circular financing” is thrown around as a pseudo intellectual gotcha, while everyone in the AI infrastructure business is witnessing demand curves that are so rapidly outpacing supply that price escalations and demand destruction are now accepted as default strategies for what is arguably the most productive and promising technology of our time. So far it has required hyperscalers to underwrite AI infrastructure at scale, but demand is larger than what hyperscaler balance sheets can securitize. It also skews the market to a small number of providers, gives them unbounded pricing leverage and defaults to their specific product strategies. With NVIDIA’s underwriting program, non-investment grade innovative startups and enterprises will be able to participate in building tomorrow’s AI infrastructure. This will result in better and more diverse products, faster AI transformation of the American industry, and a healthier and less concentrated AI infrastructure ecosystem.
Show more
This is huge news for India, Chennai, and L&T. Larsen & Toubro (L&T) has secured a major deal to build India’s largest NVIDIA B300 AI Factory. In partnership with US based AI company Together AI, L&T will set up a massive AI computing facility powered by 10,000 NVIDIA B300 GPUs. This will be India’s largest single cluster AI infrastructure of its kind. The facility will come up at Vyoma, AI’s data centre campus in Chennai and will power Together AI’s cloud platform for large scale AI training and inference. The Chennai site is built for big expansion, with Phase 1 designed for up to 250 MW of power capacity. This move significantly strengthens India’s AI infrastructure and takes the country a step closer to becoming a global hub for next generation AI. Congrats to everyone who made this happen.
Show more
0
71
2.4K
399
Forward to community
Excited to be partnering with @larsentoubro to build large AI factories in India. Our first 10K B300 cluster coming up soon in Chennai!
10,000 @nvidia B300 GPUs. India's largest AI Factory. Together AI and @larsentoubro are building the country's biggest GPU cluster, backing open-source inference, fine-tuning, and training at scale for India's AI-native ecosystem.
Show more
Very excited to partner with @IBM to build on IBM cloud and bring efficient open source AI to enterprises.
In collaboration with @togethercompute, IBM is scaling open source AI. Together AI will leverage @NVIDIA HGX B300 systems and NVIDIA Spectrum-X networking on IBM Cloud to deliver open source model inference, helping enterprises run AI workloads faster and more efficiently in production:
Show more
Some people are interpreting this as a way for OSS models to bypass encryption of closed model APIs. If that is what you are taking away from this, please understand that the title is a massive overstatement in this context. The point of encryption here is to keep the inference protocol stateless; I assume labs are fully aware of some side channel implications and do not care about preventing replay (why would they?) Encrypted traces offer a stateless inference protocol so inference backend can be distributed across hundreds or thousands of data centers without the need for coordinating sessions. If your side channel is to ask the LLM "true or false" questions for the universe of words that might be contained in a trace, you'd go broke decoding a single trace. So you'd have to assume there's some super special magic trace that lived within some encrypted blob that's worth attacking in this way. Overall definitely an interesting limitation of the encrypted vs stored traces, but you aren't going to be able to steal traces this way to train a competing model. That would cost more than building the model with the standard training arsenal.
Show more
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
Show more