가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

filipe
@filicroval
data eng | 1xAsus Ascent GX10 | benchmarking local models so you don't have to
가입 May 2023
214 팔로잉 중    153.4K
quick DeepSeek V4 Flash benchmark on one DGX Spark: → Code: 16.79 tok/s → Prose: 16.77 tok/s setup: 86.34 GiB Q2 target, ds4 runtime, greedy decode, 1,024-token context, three measured runs. increasing concurrency failed to improve aggregate tok/s in a separate run. instead, latency increased a lot i'm working on improving those numbers: profile the active bytes for every generated token, identify the tensor families dominating memory traffic, and attack the bandwidth floor every candidate gets a strict control, output checks, and an intelligence eval before it earns a speedup claim.
더 보기