๐ Congrats to
@Zai_org on GLM-5.3-Flash โ the first natively multimodal model in the GLM-5 series, and their first model combining sparse and linear attention.
GLM-5.3-Flash introduces a new combination of linear attention for local dependencies and sparse attention for global context, with IndexPool keeping the indexing overhead low even at 1M-token context.
This new architecture also brings a new set of inference challenges. TokenSpeed already has day-0 support for GLM-5.3-Flash, with the stack validated on both
@NVIDIA and
@AMD GPUs. ๐