đŖ New Inferenceing Engine Alert!
One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP.
A new engine called Athena recently dropped for Nvidia GB10 systems.
On a single 128GB DGX Spark you can get ...
DeepSeek V4 Flash
⥠8K prefill: 1,126 tps
đ 256K: 948 tps
đ Decode
@256K: 19.4 tps
Qwen3.8 Flash-Next
⥠8K prefill: 1,071 tps
đ 256K: 961 tps
đ Decode
@256K: 32.1 tps
Going from 8K â 256K barely rocks on prefill performance.
And Athena caches long conversations to disk. đđ
A 141,519-token conversation reportedly restores in:
⥠2.1 seconds vs ~2 min 20 sec to process it fresh.
It also includes
â
speculative decoding
â
OpenAI + Anthropic-compatible API
â
tool calling
â
Qwen image/document input
â
persistent agent context
â
Docker install
â
switch models without changing clients
This is what I want from DGX Spark.
đ¯ Huge models + huge context + usable speed on 1 box sitting on your desk. đ
đ Link in ALT