가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Wësche
@WescheNex1q
Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston
가입 January 2013
637 팔로잉 중    2.4K
One DGX Spark. Qwen3.6-35B 64 users. 700+ tok/s. 32,768 tokens in 54 seconds. 38W Each user has their own prompt and their own KV cache, and vLLM batches every active stream through the GPU each step. Recipe → @NVIDIAAI
더 보기