注册并分享邀请链接,可获得视频播放与邀请奖励。

Wësche
@WescheNex1q
Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston
加入 January 2013
637 正在关注    2.4K 粉丝
One DGX Spark. Qwen3.6-35B 64 users. 700+ tok/s. 32,768 tokens in 54 seconds. 38W Each user has their own prompt and their own KV cache, and vLLM batches every active stream through the GPU each step. Recipe → @NVIDIAAI
显示更多
0
64
1.1K
90
转发到社区