📣 New Inferenceing Engine Alert!
One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP.
A new engine called Athena recently dropped for Nvidia GB10 systems.
On a single 128GB DGX Spark you can get ...
DeepSeek V4 Flash
⚡ 8K prefill: 1,126 tps
📚 256K: 948 tps
🚀 Decode
@256K: 19.4 tps
Qwen3.8 Flash-Next
⚡ 8K prefill: 1,071 tps
📚 256K: 961 tps
🚀 Decode
@256K: 32.1 tps
Going from 8K → 256K barely rocks on prefill performance.
And Athena caches long conversations to disk. 👈👀
A 141,519-token conversation reportedly restores in:
⚡ 2.1 seconds vs ~2 min 20 sec to process it fresh.
It also includes
✅ speculative decoding
✅ OpenAI + Anthropic-compatible API
✅ tool calling
✅ Qwen image/document input
✅ persistent agent context
✅ Docker install
✅ switch models without changing clients
This is what I want from DGX Spark.
🎯 Huge models + huge context + usable speed on 1 box sitting on your desk. 👀
🔗 Link in ALT