註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
加入 July 2023
549 正在關注    11.2K 粉絲
Check out this part of Athena. 1 Spark can be turned into a local multi-model agent server. Install both🧠 Qwen3.8 Flash-Next & 🧠 DeepSeek V4 Flash The Spark can't hold both in its 128GB memory at once, so Athena can checkpoint the current agent/session then .. 🧹 unload one giant model 🔄 load the other ⚡ switch in ~46 sec 🔌 keep the same OpenAI/Anthropic API endpoint Your client just talks to Athena. Both quantized models scored 91/100 on Athena's tool-use eval: 🛠️ 69 scenarios 🧰 52 possible tools 🔗 chained calls 🚫 knowing when NOT to call 🔄 error recovery 📋 schema-valid JSON ⚠️ Both models had trouble with one tool-output prompt-injection scenario. Athena is free for personal/research/education use, but the engine itself is currently proprietary/closed-source.
顯示更多
📣 New Inferenceing Engine Alert! One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP. A new engine called Athena recently dropped for Nvidia GB10 systems. On a single 128GB DGX Spark you can get ... DeepSeek V4 Flash ⚡ 8K prefill: 1,126 tps 📚 256K: 948 tps 🚀 Decode @256K: 19.4 tps Qwen3.8 Flash-Next ⚡ 8K prefill: 1,071 tps 📚 256K: 961 tps 🚀 Decode @256K: 32.1 tps Going from 8K → 256K barely rocks on prefill performance. And Athena caches long conversations to disk. 👈👀 A 141,519-token conversation reportedly restores in: ⚡ 2.1 seconds vs ~2 min 20 sec to process it fresh. It also includes ✅ speculative decoding ✅ OpenAI + Anthropic-compatible API ✅ tool calling ✅ Qwen image/document input ✅ persistent agent context ✅ Docker install ✅ switch models without changing clients This is what I want from DGX Spark. 🎯 Huge models + huge context + usable speed on 1 box sitting on your desk. 👀 🔗 Link in ALT
顯示更多