here we go again: deployed a FREE public endpoint for Qwen3.8-Flash-Next 🚀 (going at +100 tok/s)
No token needed, OpenAI-compatible, vision + tool calls, 262K context, thinking from xhigh → off. Light rate limiting, be nice to your neighbors 🤗
4× H200 · FP8 · SGLang cookbook · ~140 tok/s per stream · ~100 tok/s @ 16 concurrent · 0.8s TTFT
Guide + chat UI 👇
顯示更多