Register and share your invite link to earn from video plays and referrals.

Together AI
@togethercompute
Accelerate inference, model shaping, and pre-training on a research-optimized platform.
374 Following    63.6K Followers
Our serverless APIs for open weights models have become exquisite machinery (that we plan to write more about soon!) One effect is that @togethercompute now one of the largest originators of open tokens in the US, with a wild week over week growth curve that reflects the pace of OSS adoption. Will serve 1.8T tokens today on @OpenRouter alone. You can get started with $5 —
Show more
Reliable, fast inference is essential. As Ted shared at Apsara Conference 2026, Together makes it easy for teams to get access to frontier open models without needing to manage any of the infrastructure complexity. See for yourself
Show more
Making AI simpler for businesses. At #ApsaraConference2026#, Ted Cui, VP of Engineering and Inference Platform at Together AI, spoke about AI becoming more accessible, with model selection becoming less of a concern for businesses. Find out more: #AgenticEra# #AgentNative# #AIAgentsAtApsara# #BringYourAgent# #AlibabaCloud#
Show more
Just trained Tev1 0.8B, a tiny Jev-like classifier. Here it is running completely locally on my mac with @ollama & classifying some tasks. It's extremely fast: only ~50ms E2E latency. Video is not sped up! Releasing weights & benchmarks very soon so you can try it yourself :)
Show more
"We discovered early on that it's really hard to reliably deploy AI in the real world. Frontier models are really strong, but they're really slow for real-time use cases." See how @DecagonAI partners with Together to power its voice AI concierge.
Show more
Qwen3.7-Max and Qwen3.8-Flash from @Alibaba_Qwen are now 40% off on serverless through September 30. Use Max for long-horizon agent work and Flash for high-volume, cost-sensitive workloads. Get started today!
Show more
New on Dedicated Model Inference: canary rollouts. Upgrade the model behind a live endpoint without downtime. Traffic moves from your current deployment to the new checkpoint in gated steps (default 5% → 25% → 50% → 100%). Health checks run before any traffic shifts. After every step, metric gates compare the new model's p95 latency and error rate against the old one. If a gate trips, the rollout pauses at the canary share and waits for you: resume, promote to 100%, or roll back. Three strategies: canary, blue-green, and rolling. Available now via the tg CLI, REST API, and Python SDK. Learn how to start a rollout:
Show more
Introducing Figure out which open model is best for your use case. Compare models across coding, agents, long context, vision, finance, and more. Then see how they compare on cost + quality, including what you could save by moving to open models.
Show more
coding agents are moving fast from prototype to production. the infrastructure question is what's left. join @parthsareen from @ollama and @zainhas Hasan from Together AI at @AIconference for a breakout on what it actually takes to build coding agents on open models, and run them at scale.
Show more
This week in London, we got into the real economics of running AI in production. Cost, model selection, routing, reliability. Straight talk from teams doing this at scale. Thanks to @MiniMax_AI, @deel, and everyone who joined us. Until next time.
Show more
deepseek v4.1 flash on together ai is leading on ttft for pre-warmed queries for agent workflows, that startup time shows up again on every model turn, so shaving it down matters more as the loop gets longer we’ve been pushing hard on the serving path to keep that overhead as low as possible
Show more
omg @togethercompute I love you. this is time to first token for every provider for Deepseek v4.1 flash. you guys CRUSH on pre-warmed queries. there's no other game in town
Together AI just raised $800M at an $8.3B valuation. I got their product team to open up the actual repo they run on: 0:00 - Slop is the new party foul 1:44 - Why individual output backfired 3:26 - Inside the shared product repo 8:57 - Who owns the context files 10:41 - Team skills vs personal skills 11:57 - Picking a harness and a model 15:17 - Feature research running live 16:57 - What stays human in the loop 19:56 - The PRD skill that interviews you 23:09 - Half a day of research in 5 min 24:25 - Their customer insights MCP 27:31 - What a good PRD looks like now 32:02 - Reading the finished PRD 35:07 - Every repo in one place 40:24 - Context is a hierarchy, not a pool 41:20 - Build your own orchestrator 44:19 - Agent evals on shipped features 49:41 - Design validated every few hours 52:26 - Where PM ends, engineer begins 55:10 - What it actually cost them
Show more
How do you actually shortlist an open model for production? Rochelle Mattern, our Head of Field Engineering, is tackling that at @AIconference. Real evaluation criteria, not leaderboard vibes. Worth clearing your calendar for.
Show more
Deel handles compliance, payroll, and HR questions across 150+ countries. Instead of searching for the answer through an external LLM each time, they fine-tuned small models built on Deel's own compliance expertise. Faster answers. Built on Together AI. Proud to collaborate with @deel to power the infrastructure behind their approach to global compliance.
Show more
how inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the inference engine > kv + prefix caching > continuous batching > paged attention > chunked prefill > sampling > agentic loops from inside the engine
Show more
Nice stats from @OpenRouter that illustrate how @togethercompute delivers solid and scaled performance for agentic workloads. We are serving GLM 5.3 and GLM 5.3 Flash @ top decile of TPS, latency, cache rate, and doing it at large volumes, 23% and 30% of all OpenRouter traffic, and OpenRouter is a fraction of our overall API traffic. Running scaled inference for agentic applications is a lot more than maxxing a single metric ... you need to optimize on all dimensions and doing so at scale with reliability.
Show more
This model is insane at landing pages. I asked DeepSeek V4.1 Flash & Claude Fable 5 to build me a landing page for a movie theater. Fable cost $1.21 while V4.1 Flash cost 2.6 cents, making it more than 40x cheaper at similar quality. Gave both the exact same prompt!
Show more
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
Show more
together fine-tuning now supports glm-5.3, kimi k2.7-code + more we also added live run metrics, experiment comparisons, early stopping, dataset previews + finer training controls prices are down 30–70% on selected models read more on what's new 👇
Show more
introducing preemptible compute for together gpu clusters same nvidia gpu infrastructure, 50% of the on-demand price built for evals, fine-tuning, batch inference + short experiments, with up to 5 minutes to checkpoint before reclamation now in public preview
Show more