Nice expert call on $CBRS with a Director at Figma who's sending 40% of AI traffic to Cerebras' cloud
There are clear use cases for hyperfast inference and the customer is willing to pay a premium for these fast tokens
Figma spent roughly a year building an in-house language model for its assistant product, a chat interface intended to make Figma AI-native. The model is in the 50 to 100 billion parameter range, which he characterizes as roughly five to six times smaller than the frontier models it replaces, and better in quality for the specific task. Figma ships the weights to Cerebras and Cerebras handles all serving and optimization. He explicitly frames the relationship as a cloud provider relationship, not a hardware purchase.
The AI traffic split:
- 20% of assistant traffic routes to external frontier models (Anthropic and others) for general question answering and RAG-type requests. This share is structurally permanent. Figma has no intention of training its own model for it.
- 80% routes to the in-house model.
- Of that 80%, half goes to Cerebras, so roughly 40% of total assistant traffic.
- The remaining 40% goes to Fireworks, Modal and Baseten.
Figma Make, the vibe-coding product, sits in the same org but runs almost entirely on external models from Anthropic and Google. Critically, the AI assistant is currently exposed to only around 30% of the Figma user base.
Current Cerebras spend is described as low to mid seven figures annually. His forward view: "I would expect at least in the next one year to easily double... maybe 2X-3X growth in the next one to two years."
Beyond that he expects the use case to plateau, since the core objective is penetrating the existing user base rather than acquiring new users. This is a useful anchor for how a mid-size, non-AI-native enterprise customer scales on this platform: fast doubling off a small base, then aiming to flatten once penetration completes, with the next leg dependent on new products rather than the same product growing.
# Why Cerebras Does Not Get the Other 40% of Traffic
1. Capacity is reserved, not metered. Figma has weekday peaks and effectively zero weekend and overnight traffic. Reserving for peak means paying for idle silicon most of the week, so Figma deliberately sizes its Cerebras reservation at only 40% to 50% of average traffic and sends the peaks to on-demand GPU providers. This single design choice is what creates the opening for Fireworks, Modal and Baseten.
2. Reliability - Error rates are still above what he wants from a production system, which forces Figma to maintain fallback paths.
3. Observability - He wants first-try success rates, retry counts and internal failure data, which he is not currently getting, and without which he cannot engineer around the reliability gap.
"If Cerebras did not have any limitation, we would've used the whole 80% of traffic through Cerebras."
# The Deployment Friction, Scored 7 out of 10
Model updates are not self-serve. Figma has to notify the Cerebras team roughly a week ahead. Hand a model over Monday or Tuesday, and it is deployed by Friday. Two to three days per iteration, versus effectively immediate self-serve deployment at every GPU-based competitor.
He rated the pain at seven out of ten on a scale where ten is severely damaging.
"It's of course something we can live with, but it definitely slows our execution a lot... we build a model, we do some testing, we do this three-day wait, do some testing, find that there is a small bug, and then we have to retrain the model."
This is a developer infrastructure maturity gap, and it is the kind of thing that does not show up in benchmark comparisons but does show up in renewal conversations and in how much of a customer's roadmap a vendor can capture.
# The Price of Speed
Against the frontier model, running Figma's own smaller model on Cerebras costs roughly half to one third as much for equivalent traffic. But the model is five to six times smaller, so the like-for-like inference premium is real.
Against a GPU-based provider running the same model, his estimate is that Cerebras is around 30-50% more expensive. Figma has not run a formal side-by-side, which he attributes to the fact that reserved pricing versus per-token pricing makes the answer entirely dependent on the traffic profile.
In a bake-off against hyperscalers, Modal, Baseten, Fireworks and Groq, Cerebras came out roughly 10 to 15 times faster than most of the market. Figma found Cerebras through the Artificial Analysis public benchmark, then spent months on POCs specifically to de-risk infrastructure maturity before scaling to production.
The latency threshold he describes is a genuine product constraint rather than a nice-to-have. For assistant tasks a user could perform manually in about 30 seconds, an AI response taking longer than that adds no value. Delivering in three to five seconds changes the product.
"Speed is extremely valuable... even that 50% to 2X increase in price, I think is totally something we are willing to pay for the speed."
Cost optimization at scale would likely not come from switching hardware. It would come from shrinking the model further or routing simple requests to a smaller model.
"Maybe a workflow user makes a request, it takes three hours to run. If that's the case, then speed is not a big criteria... For those use cases, we may switch to the cheaper inference provider."
Figma has not built an in-house model for Figma Make because the output format is React and HTML, which frontier models have seen extensively, leaving little quality headroom. But he notes cost and latency wins are still available, and if cost becomes a concern as Make scales, Figma would likely train a smaller model and host it on Cerebras for the base load. Additional projects beyond the assistant are early stage, with some expected to scale in the second half of the year.
He also notes Figma experimented with Groq prior to its acquisition by NVIDIA.
$CBRS $NVDA
もっと見る