So much excitement about GLM-5.3-Flash, Kimi K3, and the pace of open models in general. But many of you are asking questions such as: does it beat what's serving in prod, will the tool calls parse in my agent loop, does latency hold when the context window fills up, can I post train it on my traces, and what's the $/token at real traffic
that's what we'll cover at Builders & Brews, plus a room full of people building on open models who are worth meeting.
Find your city: