GLM-5.3-Flash is live on Fireworks on day… 2
Why? Because we take quality very seriously. We found a benchmark discrepancy we couldn’t explain, so we delayed the launch to investigate.
Day 0 (Wed): we saw 2x longer thinking on reasoning-heavy benchmarks (AIME & GPQA) for open source engines compared with
@Zai_org API. Same scores, worse token efficiency. Agentic benchmarks looked good.
We decided to investigate further, as overthinking might become a quality problem if max_tokens are reached
Day 1 (Thu): as other non-official providers launched, their APIs had thinking in the range of open-source engines: longer than
We launched a private preview endpoint with disclaimers to a few customers and worked with them to assess quality
Day 2 (Fri): the official API updates. We rerun benchmarks: reasoning is now similarly long, consistent with vllm/sglang. Rest of the benchmarks, both public and internal, check out too.
We launched GLM-5.3-Flash publicly:
More details below