가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

SGLang
@sgl_project
Run LLMs fast at any scale 🔗 Join our community For AI tech blogs & deep-dives 👉 @lmsysorg
가입 May 2025
53 팔로잉 중    9.5K 팬
Thanks to the @FireworksAI_HQ team for the patience and rigor throughout this investigation—and for sharing the findings with the community. Correctness comes first. We’re glad SGLang could be part of the collaboration to get GLM-5.3-Flash running as expected. 🚀
더 보기
GLM-5.3-Flash is live on Fireworks on day… 2 Why? Because we take quality very seriously. We found a benchmark discrepancy we couldn’t explain, so we delayed the launch to investigate. Day 0 (Wed): we saw 2x longer thinking on reasoning-heavy benchmarks (AIME & GPQA) for open source engines compared with @Zai_org API. Same scores, worse token efficiency. Agentic benchmarks looked good. We decided to investigate further, as overthinking might become a quality problem if max_tokens are reached Day 1 (Thu): as other non-official providers launched, their APIs had thinking in the range of open-source engines: longer than We launched a private preview endpoint with disclaimers to a few customers and worked with them to assess quality Day 2 (Fri): the official API updates. We rerun benchmarks: reasoning is now similarly long, consistent with vllm/sglang. Rest of the benchmarks, both public and internal, check out too. We launched GLM-5.3-Flash publicly: More details below
더 보기