Register and share your invite link to earn from video plays and referrals.

Jesse Zhang
@thejessezhang
203 Following    102.1K Followers
We're hosting our annual conference, Decagon Dialogues, in SF on 10/1! @ScottWu46 from Cognition is joining us for a keynote. We'll have folks from American Airlines, Verizon, Paypal, Chipotle, etc discussing how they're deploying agents at scale. Come by:
Show more
BCV has been awesome to work with. Congrats to the team!
Make money. Have fun. Live with integrity. Everything else can change. Fund XI. $1.6B total capital for those who know that it will.
The way engineering works is changing so quickly that you have to be very adaptable
In July, 80% of our team's AI coding happened in Claude. By August, 80% was happening in Cursor or Codex. One of the most valuable traits for engineering teams right now is a willingness to abandon a workflow they relied on just weeks ago.
Show more
Qasar Younis, CEO of @AppliedInt, has been a mentor and friend to @AshwinSreenivas since the earliest days of Decagon. He came to our SF office to share what he’s learned across two startups, Google, YC, and 8+ years building Applied Intuition. Thanks for joining us, @qasar!
Show more
Oh no they found our secret headquarters...
HUBBLE DISCOVERED THAT SATURN HAS A DECAGON AT ITS SOUTH POLE!!! SO NOW IT HAS A HEXAGON AT ITS NORTH POLE AND A DECAGON AT ITS SOUTH POLE!!
"What it comes down to is what tools you're giving it access to," says Decagon CEO @thejessezhang on enforcing guardrails on AI agents.
One cool technique to improve latency in LLMs is to have a small, fast model generate tokens and a larger, slower model review the output in one pass. It'd be like having a junior employee write a report, and then having a senior employee review it in one go before it gets sent to a client. This is called "speculative decoding" and lets you generate faster while maintaining the output quality. Great write-up from the @DecagonAI team on how we've been doing this:
Show more
We were very excited by @inco_ai's Dflash2 release and started building with it immediately. A few tweaks to the training recipe further improved acceptance length by another 33%. Full write-up below.
Show more
More great writing from our Decagon Labs team, this time on our post-training efforts!
at decagon, we’ve been building around a simple idea: the failures of your current model checkpoint should shape the training data for the next one.
What a time to be alive
fwiw I think it is _extremely_ unlikely that user data had any influence here - there is no way OAI would pull user transcripts for this, or knowingly train on it in a way that would've influenced this. I think its pretty important people don't run away with 'your user data isn't safe in codex' - because it surely is (based on everything I can assume from the outside)
Show more
Big +1 to this. We were a bit skeptical of this when we were first starting out, but once the task is defined, it's strictly better to move to smaller models. Better latency, better cost, AND you can tune it for the specific task. Great reason for frontier labs to open-source small models since they don't compete with the big models on the same use cases.
Show more
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
Show more
One interesting learning I've had is that if you have a good AI experience in place, people find new ways to interact with their customers. This usually leads to MORE interactions, not fewer. Agents should be proactive and not just reactive.
Show more
Introducing Campaign Composer, a new way to build AI agents for proactive outreach. Agents can coordinate voice and SMS across days or weeks, carry context between interactions, and follow campaign rules while working toward a customer outcome.
Show more
GPU prices have actually gone up over the last handful of months. As you scale inference aggressively, especially with open-source models, being more GPU-efficient is a big win. This is an awesome post by our Decagon Labs team about how we achieved some of these gains!
Show more
Exciting to see GLM-5.3 and GLM-5.3-Flash advance the frontier of the intelligence vs. latency tradeoff.
One of the trickiest problems you run into with voice AI in the real world is that people often call in messy environments. There might be other voices in the background, or the person might be talking to other people at the same time. Super cool work from Decagon Labs on how we've tackled this problem!
Show more
Conversational voice agents need to know when a different person is speaking, but not every background voice should count. We combined speaker embeddings with a post-trained audio-language model to determine when a speaker change matters.
Show more
Honestly impressive how good these labs are at growth hacks and drumming up hype for their models, especially given that Western social media is blocked in China...
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: Available now across all official platforms: Weights: API: Coding Plan: ZCode: Chat: AutoClaw:
Show more
This feels like the right approach. Decagon’s workflows are now roughly split 90% open source, 10% closed source. It is quite impressive how far fine-tuned open-source models can get you, and at the same time you do need a minimum amount of frontier tokens for new or novel projects. Re humans vs tokens, it's very much a positive sum situation. Having the right balance means effectively hiring a ton of output you couldn't have previously.
Show more
Suspect we will start to hear about a “Pareto optimal” balance of computationally efficient humans, cheaper open-source tokens and frontier tokens. Our internal AI spend @Atreidesmgmt will be roughly 100x higher in August 2026 vs. March 2026. Still roughly doubling every month. Note that is before Grok Bot moves to consumption pricing which will likely create a step function when it happens (at least for me). AI means that scale has become more important to investing imo. I think there will be a minimum token spend required to be competitive in most knowledge based industries. The “compute inequality” referenced by Sholto in our discussion.
Show more
Another huge milestone at @DecagonAI: We’re proud to announce our work with Delta, one of the most iconic names in global travel. As America’s most awarded airline, @Delta has built its brand on delivering a premium, reliable, and differentiated customer experience. Delta is leveraging Decagon agents to enhance the moments that matter most to customers, advancing its vision of a travel experience that feels effortless, connected, and distinctly Delta. 'Customers are everything' is a core value of ours, and it's something we share with Delta. We're proud to play a role in their continued mission to elevate the customer experience.
Show more
I've written a lot about how low latency is critical in voice agents, and it's quite a hard problem when you also need to keep accuracy super high. Here's some insight into how our team approaches a portion of this!
Show more
Fast text serving ≠ fast speech. Here's how we got our time to first audio under 30ms with nearly 10x more audio throughput.
Where you go after college has a huge impact on the rest of your life. Lots of good options, but worth thinking hard about who you surround yourself with, and hopefully it leads to a lot of learning. We love working with new grads at Decagon. Great write-up from @chenxcynthia on some of the learnings building the company so far.
Show more
Two years ago, I joined a 12 person startup as my first job out of college. A lot has changed since then. Here are a few things I learned along the way.