@ycombinator built AI versions of its partners to help more people work through their startup ideas.
For its Office Hour Simulator, the team needed useful answers delivered quickly enough for a spoken conversation. They had been testing lightweight Gemini and OpenAI models before moving to GLM-5.2 on a dedicated Wafer endpoint.
YC then compared that deployment with GPT-4.1 mini on OpenAI and Gemma 4 31B on Cerebras, evaluating answer quality, latency, and conversation duration.
The Wafer configuration delivered 31% lower average LLM latency than OpenAI and 44% lower than Cerebras. Users spent 2.5 minutes longer talking to YC’s AI partners when using Wafer.
Read how YC built the experience and landed on Wafer
🧵 link in thread