๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

wafer
@wafer_ai
Inference that Keeps Getting Better Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
๊ฐ€์ž… June 2025
10 ํŒ”๋กœ์ž‰ ์ค‘    11.5K ํŒฌ
@ycombinator built AI versions of its partners to help more people work through their startup ideas. For its Office Hour Simulator, the team needed useful answers delivered quickly enough for a spoken conversation. They had been testing lightweight Gemini and OpenAI models before moving to GLM-5.2 on a dedicated Wafer endpoint. YC then compared that deployment with GPT-4.1 mini on OpenAI and Gemma 4 31B on Cerebras, evaluating answer quality, latency, and conversation duration. The Wafer configuration delivered 31% lower average LLM latency than OpenAI and 44% lower than Cerebras. Users spent 2.5 minutes longer talking to YCโ€™s AI partners when using Wafer. Read how YC built the experience and landed on Wafer ๐Ÿงต link in thread
๋” ๋ณด๊ธฐ