Wafer (
@wafer_ai) is building a fast AI inference cloud that uses agents to optimize GPUs and run open-source models at industry-leading speeds. The result is the low cost of open source with the latency of a much smaller model.
Four months after launching the cloud, they went from zero to $8 million in ARR and just raised a $40M Series A.
In this episode of Founder Firesides, YC's
@snowmaker sits down with Wafer co-founders
@gpusteve and
@gpuemi to talk about going from UChicago undergrads to building one of the fastest-growing AI infrastructure companies in just over a year. They share how a college side project became Wafer, how they got GLM 5.2 running 2-3x faster than anyone else in the market, and why speed, not just cost, is pushing enterprises toward open-source models.
00:48 — What Wafer Does
01:35 — Going From $0 to $8M ARR in Four Months
04:21 — Scaling Faster Than They Could Keep Up
09:14 — How AI Makes AI Faster
11:21 — The Pivot That Changed Everything
14:41 — Why Open-Source Models Are Taking Off
18:03 — From a College Side Project to Wafer
23:33 — Advice for Starting and Joining Startups