๐ฃ New Inferenceing Engine Alert!
One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP.
A new engine called Athena recently dropped for Nvidia GB10 systems.
On a single 128GB DGX Spark you can get ...
DeepSeek V4 Flash
โก 8K prefill: 1,126 tps
๐ 256K: 948 tps
๐ Decode
@256K: 19.4 tps
Qwen3.8 Flash-Next
โก 8K prefill: 1,071 tps
๐ 256K: 961 tps
๐ Decode
@256K: 32.1 tps
Going from 8K โ 256K barely rocks on prefill performance.
And Athena caches long conversations to disk. ๐๐
A 141,519-token conversation reportedly restores in:
โก 2.1 seconds vs ~2 min 20 sec to process it fresh.
It also includes
โ
speculative decoding
โ
OpenAI + Anthropic-compatible API
โ
tool calling
โ
Qwen image/document input
โ
persistent agent context
โ
Docker install
โ
switch models without changing clients
This is what I want from DGX Spark.
๐ฏ Huge models + huge context + usable speed on 1 box sitting on your desk. ๐
๐ Link in ALT