๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
Check out this part of Athena. 1 Spark can be turned into a local multi-model agent server. Install both๐Ÿง  Qwen3.8 Flash-Next & ๐Ÿง  DeepSeek V4 Flash The Spark can't hold both in its 128GB memory at once, so Athena can checkpoint the current agent/session then .. ๐Ÿงน unload one giant model ๐Ÿ”„ load the other โšก switch in ~46 sec ๐Ÿ”Œ keep the same OpenAI/Anthropic API endpoint Your client just talks to Athena. Both quantized models scored 91/100 on Athena's tool-use eval: ๐Ÿ› ๏ธ 69 scenarios ๐Ÿงฐ 52 possible tools ๐Ÿ”— chained calls ๐Ÿšซ knowing when NOT to call ๐Ÿ”„ error recovery ๐Ÿ“‹ schema-valid JSON โš ๏ธ Both models had trouble with one tool-output prompt-injection scenario. Athena is free for personal/research/education use, but the engine itself is currently proprietary/closed-source.
๋” ๋ณด๊ธฐ
๐Ÿ“ฃ New Inferenceing Engine Alert! One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP. A new engine called Athena recently dropped for Nvidia GB10 systems. On a single 128GB DGX Spark you can get ... DeepSeek V4 Flash โšก 8K prefill: 1,126 tps ๐Ÿ“š 256K: 948 tps ๐Ÿš€ Decode @256K: 19.4 tps Qwen3.8 Flash-Next โšก 8K prefill: 1,071 tps ๐Ÿ“š 256K: 961 tps ๐Ÿš€ Decode @256K: 32.1 tps Going from 8K โ†’ 256K barely rocks on prefill performance. And Athena caches long conversations to disk. ๐Ÿ‘ˆ๐Ÿ‘€ A 141,519-token conversation reportedly restores in: โšก 2.1 seconds vs ~2 min 20 sec to process it fresh. It also includes โœ… speculative decoding โœ… OpenAI + Anthropic-compatible API โœ… tool calling โœ… Qwen image/document input โœ… persistent agent context โœ… Docker install โœ… switch models without changing clients This is what I want from DGX Spark. ๐ŸŽฏ Huge models + huge context + usable speed on 1 box sitting on your desk. ๐Ÿ‘€ ๐Ÿ”— Link in ALT
๋” ๋ณด๊ธฐ