quick DeepSeek V4 Flash benchmark on one DGX Spark:
→ Code: 16.79 tok/s
→ Prose: 16.77 tok/s
setup: 86.34 GiB Q2 target, ds4 runtime, greedy decode, 1,024-token context, three measured runs.
increasing concurrency failed to improve aggregate tok/s in a separate run. instead, latency increased a lot
i'm working on improving those numbers: profile the active bytes for every generated token, identify the tensor families dominating memory traffic, and attack the bandwidth floor
every candidate gets a strict control, output checks, and an intelligence eval before it earns a speedup claim.