Register and share your invite link to earn from video plays and referrals.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
549 Following    11.2K Followers
You have to check out this Jev Repro-Clone tracker on HuggingFace. link in ALT.
Did you know you can now put the text of almost all of arXiv on one cheap SSD. Someone just packaged an enormous arXiv snapshot on Hugging Face ๐Ÿ“š 3,148,796 papers ๐Ÿ“„ PDFs for 99.47% ๐Ÿ’ฝ Full archive: 16.08 TB Sounds ridiculous for Local AIโ€ฆ Until you see that โ€ฆ ๐Ÿ”ฅ paper_text = 70GB That contains the resolved TeX text for โ€ฆ ๐Ÿง  2,856,227 papers ๐Ÿ“Š 90.7% of all arXiv And you can download just that dataset by itself. No 16TB NAS is necessary. So theoretically you could build a completely local research system using ThumbLLM.exe from a thumb drive!!! ๐Ÿ’พ 70GB arXiv text corpus ๐Ÿ”Ž local search / embeddings ๐Ÿง  local LLM ๐Ÿšซ no API required Ask your local model to search nearly 3 million scientific papers sitting on your own SSD. ๐Ÿ‘€ Caveats โš ๏ธ Itโ€™s a snapshot, not live arXiv โš ๏ธ latest submission is Aug. 27 โš ๏ธ the text is resolved TeX, so it still contains LaTeX syntax/macros and needs some cleaning for ideal RAG But 2.8M papers in ~70GB is a Localmaxxer dataset. Link in ALT
Show more
K2 Horizon got another llama.cpp-family path, this time through TurboQuant. Spark-X2.5 already landed in upstream llama.cpp. Iโ€™ve actually got it running myself. A brand-new K2 Horizon TurboQuant PR adds โ€ฆ ๐Ÿง  K2 Horizon model support ๐Ÿ“ฆ Hugging Face โ†’ GGUF conversion ๐Ÿ’ฌ tokenizer + chat templates โš™๏ธ llama.cpp-style inference And theyโ€™ve already tested ๐Ÿง  K2-Horizon-3.7B Q4_K_M ๐ŸŽ M3 Pro / 18GB โšก ~35 tok/s ๐Ÿ“š 64K configured context The PR also adds Spark-X2.5 to the TurboQuant fork: โšก ~36 tps ๐Ÿ“š 256K configured context Caveats โš ๏ธ This is still an OPEN PR โš ๏ธ Only Metal has been tested โš ๏ธ Other K2 sizes + MoVA/MoE variants remain untested K2 Horizon already has an IFM llama.cpp fork, but it still isnโ€™t supported in normal upstream llama.cpp. So what Iโ€™m watching here is whether TurboQuant becomes another path for K2.
Show more
How did I miss this?! Bonsai isnโ€™t just an LLM family. Back in May, PrismML released Bonsai Image 4B, a crazy low-bit version of FLUX.2 Klein 4B designed to run locally on ๐ŸŽ iPhones. Look at this ๐ŸŽจ FLUX.2 Klein 4B transformer ๐Ÿ’พ FP16: 7.75GB ๐ŸŒณ Ternary Bonsai: 1.21GB And the entire Apple deployment payload, including its compressed text encoder + VAE, is only: ๐Ÿ”ฅ 3.88GB ๐Ÿ“ฑ iPhone 17 Pro Max ๐Ÿง  A19 Pro / 12GB unified memory ๐Ÿ–ผ๏ธ 512ร—512 โšก 9.4 sec/image ๐Ÿ“ฒ MLX Swift โ˜๏ธ No cloud ๐Ÿ’พ Bomsai Studio (App Store) M4 Pro: ~5.8 sec/image. No cloud. The ~4B diffusion transformer is paired with a 4-bit Qwen3-4B text encoder, which gets unloaded after the prompt is encoded to save memory. And thereโ€™s also ๐ŸŽ Mac / iPhone / iPad support ๐ŸŸข low-bit Gemlite builds for Nvidia GPUs ๐Ÿ”“ Apache 2.0 This is completely separate from the Bonsai 2 27B LLM Iโ€™ve been posting about. PrismML took a 7.75GB FLUX transformer and crunched it down to 1.21G and put image generation on an iPhone. ๐Ÿ‘€ And somehow I missed this for 4 months.
Show more
๐Ÿคฏ A 600B parameter model is going open-weight Oct. 15โ€ฆ โ€ฆbut only 27B parameters are active per token. ๐Ÿ‘ˆ๐Ÿ‘€ Here is Step 5 Preview (this model is exciting) and here is why โ€ฆ Stats ๐Ÿ‘‡ ๐Ÿง  600B total parameters โšก 27B active/token ๐Ÿ“š 1M context ๐Ÿ‘๏ธ vision + video ๐Ÿค– built for long-running agents ๐Ÿ”“ open weights Oct. 15 And that 27B-active number makes this really interesting for Local AI right? At roughly 4-bit, 600B parameters would theoretically be ~300GB of weights. Real-world quantization + overhead means Iโ€™d expect something more like ~320โ€“350GB before accounting for KV cache and other runtime memory. So weโ€™re potentially looking at ๐Ÿ’พ ~384GB-class memory โ†’ Q4 territory ๐Ÿ’พ ~256GB-class memory โ†’ aggressive Q3 territory And because this is MoE, only 27B parameters participate in each tokenโ€™s compute. โš ๏ธ That does NOT mean this is a 27B model or that itโ€™ll run in 27B-sized VRAM. The entire model still has to live somewhere. But with the right runtime, that could mean: Tiered memory !!!! ๐ŸŽฎ GPU โ†’ active compute ๐Ÿง  RAM โ†’ expert weights ๐Ÿ’ฝ SSD โ†’ colder experts/offload And hereโ€™s the comparison that I really like Kimi K3: ~2.8T/ ~104B active Step 5: 600B / 27B active Yet both currently land at 44 on Artificial Analysisโ€™ Intelligence Index. So Step 5 may deliver Kimi K3-class intelligence with roughly: ๐Ÿ”ฅ 79% fewer total parameters ๐Ÿ”ฅ 74% fewer active parameters This could make Step 5 a much more realistic monster model for local hardware. When the weights drop Oct. 15, my first question how small a machine can we get this thing running on? ๐Ÿ‘€
Show more
Qwenโ€™s Image 2.1 released but it come with lots of improvements but just as many caveats. The old Qwen-Image-2512 was a 20B model with a massive: ๐Ÿ’พ 40.9GB BF16 image transformer But the new Qwen-Image-2.1 comes with โ€ฆ ๐Ÿง  7B visual generator ๐Ÿ’พ 14.2GB BF16 ๐Ÿ”ฅ 7.26GB INT8 already available for ComfyUI ๐Ÿ‘ˆ๐Ÿ‘€ (Day-1) This little thing can โ€ฆ ๐ŸŽจ generate AND edit images ๐Ÿ–ผ๏ธ generate native 2K ๐Ÿซฅ create real transparent RGBA images ๐Ÿ‘ฅ use up to 10 reference images โœ๏ธ preserve people/products while editing ๐Ÿ”ค render text ๐ŸŽฏ do masked/local edits Qwen basically took the huge local image model and reduced with a catch โ€ฆ โš ๏ธ The tiny 7.26GB image model comes with an entire pipeline. It uses a Qwen3-VL 8B encoder + VAE. ComfyUI already has a 6.31GB W4A8 encoder, and Qwen supports CPU offloading for smaller GPUs. Alsoโ€ฆ ๐Ÿšซ Qwen-Image-2.1 is NON-COMMERCIAL under the new Qwen Research License. Thatโ€™s especially notable because Qwen-Image-2512 was Apache 2.0.
Show more
On the SAME 48GB M5 Pro running ๐Ÿง  Qwen3.6-35B-A3B @ 210 tps and @ 143 tok/s at 32K context, the multi-agent handling is so much better than mere tokens. With 32K of context already cached, Splash produced the first token in โšก 123 ms vs ๐ŸŒ 816 ms for oMLX And when Inco tested 16 simultaneous 32K requests with Qwen3.8-27Bโ€ฆ Splash handled all 16. Itโ€™s built to handle heavy agent workloads with lots of context and repeated prompts even with multiple agents at once w/fast reuse of cached context
Show more
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source inference engine called Splash. Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model ๐Ÿ‘‡ โš™๏ธ model-specific kernels ๐Ÿง  hardware-aware memory planning ๐Ÿš€ DFlash2 speculative decoding ๐Ÿ’พ prompt-cache reuse ๐Ÿ‘ฅ continuous batching On the same 48GB M5 Pro running Qwen3.8-27B: ๐Ÿš€ Splash: 74 tok/s โšก oMLX: 38 tok/s ๐ŸŒ Ollama: 24 tok/s At 32K context: ๐Ÿš€ Splash: 54 tok/s And with 4 concurrent requests: ๐Ÿ”ฅ 170 aggregate tok/s vs 43 tok/s for oMLX Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max. And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime. Requirements ๐ŸŽ M3 or newer ๐Ÿ’พ 36GB minimum ๐Ÿ‘ 48GB+ recommended Completely different performance because the software stack is optimized around what it is running. โš ๏ธ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test. ๐Ÿ”—
Show more
Check out this part of Athena. 1 Spark can be turned into a local multi-model agent server. Install both๐Ÿง  Qwen3.8 Flash-Next & ๐Ÿง  DeepSeek V4 Flash The Spark can't hold both in its 128GB memory at once, so Athena can checkpoint the current agent/session then .. ๐Ÿงน unload one giant model ๐Ÿ”„ load the other โšก switch in ~46 sec ๐Ÿ”Œ keep the same OpenAI/Anthropic API endpoint Your client just talks to Athena. Both quantized models scored 91/100 on Athena's tool-use eval: ๐Ÿ› ๏ธ 69 scenarios ๐Ÿงฐ 52 possible tools ๐Ÿ”— chained calls ๐Ÿšซ knowing when NOT to call ๐Ÿ”„ error recovery ๐Ÿ“‹ schema-valid JSON โš ๏ธ Both models had trouble with one tool-output prompt-injection scenario. Athena is free for personal/research/education use, but the engine itself is currently proprietary/closed-source.
Show more
๐Ÿ“ฃ New Inferenceing Engine Alert! One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP. A new engine called Athena recently dropped for Nvidia GB10 systems. On a single 128GB DGX Spark you can get ... DeepSeek V4 Flash โšก 8K prefill: 1,126 tps ๐Ÿ“š 256K: 948 tps ๐Ÿš€ Decode @256K: 19.4 tps Qwen3.8 Flash-Next โšก 8K prefill: 1,071 tps ๐Ÿ“š 256K: 961 tps ๐Ÿš€ Decode @256K: 32.1 tps Going from 8K โ†’ 256K barely rocks on prefill performance. And Athena caches long conversations to disk. ๐Ÿ‘ˆ๐Ÿ‘€ A 141,519-token conversation reportedly restores in: โšก 2.1 seconds vs ~2 min 20 sec to process it fresh. It also includes โœ… speculative decoding โœ… OpenAI + Anthropic-compatible API โœ… tool calling โœ… Qwen image/document input โœ… persistent agent context โœ… Docker install โœ… switch models without changing clients This is what I want from DGX Spark. ๐ŸŽฏ Huge models + huge context + usable speed on 1 box sitting on your desk. ๐Ÿ‘€ ๐Ÿ”— Link in ALT
Show more
The more I look at setups like this, the more I think OCuLink changes the perspective on Local AI mini-PCs. That old RTX 2080 Super doesn't need to hold your giant LLM to be useful. Give the NVIDIA GPU the CUDA-friendly jobs like ๐Ÿ‘๏ธ vision models and ๐ŸŽจ image generation and ... mini-PC's CPU/iGPU/system RAM handles the larger model or another workload when 8GB isn't enough.
Show more
Loving this OCuLink setup with the GMKtec EVO-T1 paired with an RTX 2080 Super. An Intel mini-PC gets dedicated Nvidia CUDA + 8GB of VRAM over PCIe 4.0 ร—4 OCuLink. Local AI possibilities are endless! Could run separate models/workloads across the iGPU/CPU and NVIDIA GPU or experiment with model offloading, or even play with disaggregated inference!
Show more
๐Ÿ“ฃ New Inferenceing Engine Alert! One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP. A new engine called Athena recently dropped for Nvidia GB10 systems. On a single 128GB DGX Spark you can get ... DeepSeek V4 Flash โšก 8K prefill: 1,126 tps ๐Ÿ“š 256K: 948 tps ๐Ÿš€ Decode @256K: 19.4 tps Qwen3.8 Flash-Next โšก 8K prefill: 1,071 tps ๐Ÿ“š 256K: 961 tps ๐Ÿš€ Decode @256K: 32.1 tps Going from 8K โ†’ 256K barely rocks on prefill performance. And Athena caches long conversations to disk. ๐Ÿ‘ˆ๐Ÿ‘€ A 141,519-token conversation reportedly restores in: โšก 2.1 seconds vs ~2 min 20 sec to process it fresh. It also includes โœ… speculative decoding โœ… OpenAI + Anthropic-compatible API โœ… tool calling โœ… Qwen image/document input โœ… persistent agent context โœ… Docker install โœ… switch models without changing clients This is what I want from DGX Spark. ๐ŸŽฏ Huge models + huge context + usable speed on 1 box sitting on your desk. ๐Ÿ‘€ ๐Ÿ”— Link in ALT
Show more
Interesting HuggingFace post. What if the easiest way to make long-context Local AI faster is to simply STOP feeding it everything? A new HF forum experiment tested 600 distractor-heavy long-context prompts with Qwen2.5-7B 4-bit. Instead of sending the entire ~14K-token context, a MiniLM semantic selector kept only the most relevant ~60%. The result ... ๐Ÿ“š Input tokens 13,124 โ†’ 7,702 โžก๏ธ -41% โšก Prefill 871 ms โ†’ 463 ms โžก๏ธ -47% ๐Ÿš€ End-to-end latency 944 ms โ†’ 519 ms โžก๏ธ -45% ๐Ÿ’พ Peak VRAM 11.71 GiB โ†’ 9.16 GiB And here's the weird part... ๐ŸŽฏ Token F1 actually IMPROVED: 0.534 โ†’ 0.581 So the model got: โœ… less context โœ… lower VRAM use โœ… almost 2ร— faster prefill โœ… lower total latency โœ… BETTER answers Why? Because more context isn't always better context. Coding agents and RAG systems often drag around enormous histories filled with irrelevant junk. โš ๏ธ This was a controlled distractor-heavy HotpotQA stress test, not proof that deleting 40% of every real-world context will improve accuracy. ๐Ÿ”— Link in ALT.
Show more
Intel's 32GB Arc Pro B70 finally got a proper Qwen3.8-27B Local AI workout. So this is why the 32GB VRAM is important, even from an Intel GPU. Here is Gigazine's setup. Windows 11 using Unsloth Desktop + llama.cpp ... ๐Ÿง  Qwen3.8-27B Q4_K_XL ๐Ÿ’พ 30.0GB VRAM + ~1GB shared RAM ๐Ÿš€ 23.4 tps Then they tried Q8: ๐Ÿง  Qwen3.8-27B Q8_K_XL ๐Ÿ’พ 30.2GB VRAM + ~1.5GB shared RAM โšก 15.9 tps That's a full 27B-class model running at very usable speeds on an Intel GPU. And also tested text-to-image models. ๐ŸŽจ Z-Image-Turbo 1024ร—1024 / 8 steps โžก๏ธ ~5.3 sec/image after warmup ๐ŸŽฅ MiniMax H3 video also ran โ€ฆbut these are much slower, demonstrating where Nvidia's more mature AI software stack still matters. The hardware setup... ๐ŸŽฎ Arc Pro B70 ๐Ÿ’พ 32GB GDDR6 โšก 608 GB/s bandwidth ๐Ÿ”Œ 230W And here's the interesting value angle. In Japan, the tested ASRock B70 was selling for: ๐Ÿ’ฐ ยฅ298,054 while many RTX 5090s were selling above: ๐Ÿ’ฐ ยฅ900,000 โš ๏ธ That's Japan-specific street pricing, NOT a universal B70-vs-5090 price comparison. But THIS is why the B70 interests me for Local AI. It won't beat a 5090 but it's giving you 32GB of VRAM at a much lower entry point. 23.4 tok/s on Qwen3.8-27B Q4 good enough for you? ๐Ÿ”— GIGAZINE review in ALT
Show more
Okay, so specialized decision models are becoming the rage. Bespoke Labs just released Nimble, an open 9B decision model built from Qwen3.5-9B. And apparently they built it in ONE DAY. The recipe: ๐Ÿง  Qwen3.5-9B ๐ŸŽฏ LoRA fine-tune ๐Ÿ“š only 2,676 training examples โŒ no Jev distillation โŒ no RL โœ… open model โœ… open data โœ… open training recipe Instead of generating explanations or JSON, Nimble scores the allowed answers: YES / NO โ†’ A / B / C โ†’ route 1 / 2 / 3 โ†’ severity 1โ€“5 Then returns the choice + probabilities. Example of Bespoke's 324-example held-out test: Qwen3.5-9B โ†’ 66.36% Qwen3.8-27B โ†’ 84.88% ๐Ÿ”ฅ Nimble-9B โ†’ 90.12% Jev โ†’ 93.21% And because it isn't generating a bunch of tokens: โšก ~106 ms median on H100 ๐ŸŽ ~444 ms median on an M5 Pro 64GB ๐Ÿ‘ˆ ๐Ÿ‘€ ๐ŸŽฏ So this is the key, giant reasoning model doesn't need to answer every tiny agent decision. Let the big model think, then let a small local model handle the ... ๐Ÿ› ๏ธ tool selection ๐Ÿงญ routing โœ… verification ๐Ÿšจ policy decisions ๐Ÿ“Š scoring ๐Ÿ”„ retry / stop Fast. Local. No API call. โš ๏ธ Bespoke says this is a narrow synthetic 324-example evaluation, NOT a standardized System One benchmark. ๐Ÿ”— GH: /bespokelabsai/nimble
Show more
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source inference engine called Splash. Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model ๐Ÿ‘‡ โš™๏ธ model-specific kernels ๐Ÿง  hardware-aware memory planning ๐Ÿš€ DFlash2 speculative decoding ๐Ÿ’พ prompt-cache reuse ๐Ÿ‘ฅ continuous batching On the same 48GB M5 Pro running Qwen3.8-27B: ๐Ÿš€ Splash: 74 tok/s โšก oMLX: 38 tok/s ๐ŸŒ Ollama: 24 tok/s At 32K context: ๐Ÿš€ Splash: 54 tok/s And with 4 concurrent requests: ๐Ÿ”ฅ 170 aggregate tok/s vs 43 tok/s for oMLX Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max. And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime. Requirements ๐ŸŽ M3 or newer ๐Ÿ’พ 36GB minimum ๐Ÿ‘ 48GB+ recommended Completely different performance because the software stack is optimized around what it is running. โš ๏ธ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test. ๐Ÿ”—
Show more
Nvidia dropped an official DeepSeek-V4.1-Flash NVFP4 build. And this thing is BIG. DeepSeek V4.1 Flash stats ๐Ÿง  552B backbone ๐Ÿ“š +196B Engram conditional memory โšก only 8B active during prefill ๐Ÿš€ 16B active during decode ๐Ÿ‘๏ธ native vision ๐Ÿ“– 1 MILLION token context ๐Ÿงฉ 384 routed experts across 40 layers ๐Ÿ“œ MIT license Nvidia has converted its routed MoE experts to NVFP4 W4A4 specifically for Blackwell GPUs. This is not a compressed giant model down to 4-bit. DeepSeek's experts were already stored in MXFP4, Nvidia instead converts them to its Blackwell-friendly NVFP4 format. And because NVFP4 uses finer scaling, the checkpoint actually gets slightly larger. ๐Ÿ’พ Source: ~476 GiB ๐Ÿ’พ NVIDIA NVFP4: ~492 GiB 48 safetensor shards. ๐Ÿ˜ณ So why bother? Because NVIDIA is optimizing how those 4-bit experts execute on Blackwell. And impressively, NVIDIA's evaluations show basically no obvious quality collapse from the conversion. For example: ๐Ÿง  GPQA Diamond 91.04 โ†’ 91.29 ๐Ÿ’ป SciCode 54.40 โ†’ 55.84 ๐Ÿ› ๏ธ Terminal-Bench 2.1 81.60 โ†’ 82.16 ๐Ÿ‘๏ธ MMMU-Pro 74.05 โ†’ 73.70 Some slightly up. Some slightly down. Essentially benchmark parity. And it already has: โœ… vLLM support โœ… SGLang support โœ… reasoning parser โœ… tool calling โœ… image input โœ… 1M context โš ๏ธ Nvidia validated it on 4ร— GB300 GPUs so not a local model (yet). The checkpoint is still ~492 GiB. ๐Ÿ”— HF: /nvidia/DeepSeek-V4.1-Flash-NVFP4
Show more
๐Ÿคฏ This is like having Jev at home, a AI reflex system running locally. Someone is turning Google's open DiffusionGemma 26B-A4B into a local version of the new โ€œSystem Oneโ€ AI idea inside vLLM. Remember Jev? It is a decision LLM. Instead of asking an LLM to generate paragraphs, System One models are designed to make FAST structured decisions: ๐ŸŽฏ yes / no ๐Ÿงญ route A / B / C ๐Ÿ› ๏ธ which tool to call ๐Ÿšจ severity 1โ€“4 ๐Ÿ“Š classify / score / choose And DiffusionGemma has an important role because a normal autoregressive LLM takes Prompt โ†“ token โ†“ answer DiffusionGemma is Prompt โ†“ predefined answer slots โ†“ fill multiple decisions in parallel Because it operates over a token canvas with bidirectional attention instead of being forced to generate everything strictly left-to-right. The vLLM patch lets it return โœ… bounded choices โœ… probabilities โœ… confidence / uncertainty โœ… multiple decisions simultaneously And it runs locally. On ONE DGX Spark the developer reports โšก 1 request: ~120 ms ๐Ÿš€ concurrency 32: ~54 requests/sec ๐Ÿง  3 decisions/request ๐Ÿ”ฅ ~162 decisions/sec Then the developer found each question can require only around 3 canvas tokens: ๐Ÿ”ข index โ“ placeholder โœ‚๏ธ separator Meaning as many as ~85 questions could fit into 1 diffusion canvas. And subsequent optimization reportedly cut decision time another: โšก ~20โ€“40% with no change in decision quality. The model stats ... ๐Ÿง  25.2B total parameters โšก 3.8B active ๐Ÿ’พ NVIDIA NVFP4: ~18.9GB ๐ŸŽฎ Can fit on a 24GB NVIDIA GPU ๐Ÿ“œ Open weights โš ๏ธ vLLM PR #57250# is still OPEN ๐Ÿ‘‰ not merged. And this is Jev-like functionality. It is NOT evidence that DiffusionGemma matches Jev's intelligence or calibration. ๐Ÿ”— /vllm-project/vllm/pull/57250
Show more
Loving this OCuLink setup with the GMKtec EVO-T1 paired with an RTX 2080 Super. An Intel mini-PC gets dedicated Nvidia CUDA + 8GB of VRAM over PCIe 4.0 ร—4 OCuLink. Local AI possibilities are endless! Could run separate models/workloads across the iGPU/CPU and NVIDIA GPU or experiment with model offloading, or even play with disaggregated inference!
Show more
Been using OpenAI's new Image Gen 2.5. Here is one of my favorites generated with the new tool.
๐Ÿคฏ Remember back in July I called Ling-3.0-flash one of the more interesting Local AI models? ... because Ling-3.0-flash combined 124B-scale capacity with only ~5B active parameters per token, strong frontierish benchmarks, and a design that looked promising for local agents. My July post ended with one BIG problem ๐Ÿ‘‰ localmaxxers couldn't download the weights or load it into llama.cpp.โ€ Well... that part changed. ๐Ÿ”ฅ @TheInclusionAI has now released the weights under MIT, an official GGUF exists for Ling-3.0-flash, and they've just released Ling-3.0-flash-VL. The new VL model adds: ๐Ÿ‘€ Images ๐ŸŽฅ Video ๐Ÿค– Visual/computer-use agents ๐Ÿง  124B total / ~5.5B active ๐Ÿ“š up to 1M context โš–๏ธ MIT A community Ling-3.0-flash-VL Q4_K_M GGUF is already ~79.3GB ~1.7GB vision projector. So roughly 81GB for a 124B multimodal MoE. That puts it in range of these devices ... ๐Ÿง  128GB Strix Halo ๐ŸŽ high-memory Apple Silicon ๐ŸŽฎ large/multi-GPU home rigs โš ๏ธ The VL GGUF is still experimental and currently uses a patched llama.cpp, so this is not yet a clean stock-llama.cpp experience. But this is exactly the update I wanted back in July, weights on our own machines. ๐Ÿ”ฅ Now somebody run that ~81GB VL stack on a 128GB Strix Halo and give us the tps. ๐Ÿ‘€ ๐Ÿ”— HF inclusionAI/Ling-3.0-flash-VL
Show more
๐Ÿคฏ Remember the weird U.S.-built diffusion LLM I posted about that broke 1,000 tok/s on Nvidia GPUs? It got a BIG upgrade. @_inception_ai released Mercury 2.5 today. ๐ŸŽฏ Instead of generating one token after another like a normal LLM, Mercury uses diffusion to generate/refine multiple tokens in parallel. And Inception says the new model delivers the following ๐Ÿ‘‡ ๐Ÿš€ 1,107 tok/s ๐Ÿ‘€ ๐Ÿง  +40% intelligence vs Mercury 2 ๐Ÿ“š 260K context ๐Ÿค” Tunable reasoning ๐Ÿ”ง Parallel tool calling ๐Ÿ“ Structured JSON ๐Ÿ’ต $0.20/M input / $0.75/M output ๐Ÿ‘€ ๐ŸŽฏInception says this is the largest diffusion language model ever trained. That ~1,100 tok/s isn't coming from Groq or Cerebras-style custom inference hardware - this is an important point! It's running on widely available NVIDIA GPU infrastructure. That's the big architectural bet, traditional LLM ๐Ÿ‘‡ token โ†’ token โ†’ token โ†’ token Diffusion LLM ๐Ÿ‘‡ many tokens โ†’ refine them together โ†’ answer โš ๏ธ Caveat, of course! ๐Ÿ”’ Mercury 2.5 is closed-weights.
Show more
Today, weโ€™re introducing Mercury 2.5, the most capable diffusion LLM on the market. It offers a 40% jump in intelligence over Mercury 2, and runs over 1,100 tokens/sec on widely-available @NVIDIAAI GPUs. Itโ€™s available today on our API, @OpenRouter, and @Baseten. Contact us to evaluate Mercury 2.5 for production:
Show more