You have to check out this Jev Repro-Clone tracker on HuggingFace. link in ALT.
Did you know you can now put the text of almost all of arXiv on one cheap SSD.
Someone just packaged an enormous arXiv snapshot on Hugging Face
๐ 3,148,796 papers
๐ PDFs for 99.47%
๐ฝ Full archive: 16.08 TB
Sounds ridiculous for Local AIโฆ
Until you see that โฆ
๐ฅ paper_text = 70GB
That contains the resolved TeX text for โฆ
๐ง 2,856,227 papers
๐ 90.7% of all arXiv
And you can download just that dataset by itself.
No 16TB NAS is necessary.
So theoretically you could build a completely local research system using ThumbLLM.exe from a thumb drive!!!
๐พ 70GB arXiv text corpus
๐ local search / embeddings
๐ง local LLM
๐ซ no API required
Ask your local model to search nearly 3 million scientific papers sitting on your own SSD. ๐
Caveats
โ ๏ธ Itโs a snapshot, not live arXiv
โ ๏ธ latest submission is Aug. 27
โ ๏ธ the text is resolved TeX, so it still contains LaTeX syntax/macros and needs some cleaning for ideal RAG
But 2.8M papers in ~70GB is a Localmaxxer dataset.
Link in ALT
Show more
K2 Horizon got another llama.cpp-family path, this time through TurboQuant.
Spark-X2.5 already landed in upstream llama.cpp. Iโve actually got it running myself.
A brand-new K2 Horizon TurboQuant PR adds โฆ
๐ง K2 Horizon model support
๐ฆ Hugging Face โ GGUF conversion
๐ฌ tokenizer + chat templates
โ๏ธ llama.cpp-style inference
And theyโve already tested
๐ง K2-Horizon-3.7B Q4_K_M
๐ M3 Pro / 18GB
โก ~35 tok/s
๐ 64K configured context
The PR also adds Spark-X2.5 to the TurboQuant fork:
โก ~36 tps
๐ 256K configured context
Caveats
โ ๏ธ This is still an OPEN PR
โ ๏ธ Only Metal has been tested
โ ๏ธ Other K2 sizes + MoVA/MoE variants remain untested
K2 Horizon already has an IFM llama.cpp fork, but it still isnโt supported in normal upstream llama.cpp.
So what Iโm watching here is whether TurboQuant becomes another path for K2.
Show more
How did I miss this?! Bonsai isnโt just an LLM family.
Back in May, PrismML released Bonsai Image 4B, a crazy low-bit version of FLUX.2 Klein 4B designed to run locally on ๐ iPhones.
Look at this
๐จ FLUX.2 Klein 4B transformer
๐พ FP16: 7.75GB
๐ณ Ternary Bonsai: 1.21GB
And the entire Apple deployment payload, including its compressed text encoder + VAE, is only:
๐ฅ 3.88GB
๐ฑ iPhone 17 Pro Max
๐ง A19 Pro / 12GB unified memory
๐ผ๏ธ 512ร512
โก 9.4 sec/image
๐ฒ MLX Swift
โ๏ธ No cloud
๐พ Bomsai Studio (App Store)
M4 Pro: ~5.8 sec/image.
No cloud.
The ~4B diffusion transformer is paired with a 4-bit Qwen3-4B text encoder, which gets unloaded after the prompt is encoded to save memory.
And thereโs also
๐ Mac / iPhone / iPad support
๐ข low-bit Gemlite builds for Nvidia GPUs
๐ Apache 2.0
This is completely separate from the Bonsai 2 27B LLM Iโve been posting about.
PrismML took a 7.75GB FLUX transformer and crunched it down to 1.21G and put image generation on an iPhone. ๐
And somehow I missed this for 4 months.
Show more
๐คฏ A 600B parameter model is going open-weight Oct. 15โฆ
โฆbut only 27B parameters are active per token. ๐๐
Here is Step 5 Preview (this model is exciting) and here is why โฆ
Stats ๐
๐ง 600B total parameters
โก 27B active/token
๐ 1M context
๐๏ธ vision + video
๐ค built for long-running agents
๐ open weights Oct. 15
And that 27B-active number makes this really interesting for Local AI right?
At roughly 4-bit, 600B parameters would theoretically be ~300GB of weights.
Real-world quantization + overhead means Iโd expect something more like ~320โ350GB before accounting for KV cache and other runtime memory.
So weโre potentially looking at
๐พ ~384GB-class memory โ Q4 territory
๐พ ~256GB-class memory โ aggressive Q3 territory
And because this is MoE, only 27B parameters participate in each tokenโs compute.
โ ๏ธ That does NOT mean this is a 27B model or that itโll run in 27B-sized VRAM.
The entire model still has to live somewhere.
But with the right runtime, that could mean:
Tiered memory !!!!
๐ฎ GPU โ active compute
๐ง RAM โ expert weights
๐ฝ SSD โ colder experts/offload
And hereโs the comparison that I really like
Kimi K3: ~2.8T/ ~104B active
Step 5: 600B / 27B active
Yet both currently land at 44 on Artificial Analysisโ Intelligence Index.
So Step 5 may deliver Kimi K3-class intelligence with roughly:
๐ฅ 79% fewer total parameters
๐ฅ 74% fewer active parameters
This could make Step 5 a much more realistic monster model for local hardware.
When the weights drop Oct. 15, my first question how small a machine can we get this thing running on? ๐
Show more
Qwenโs Image 2.1 released but it come with lots of improvements but just as many caveats.
The old Qwen-Image-2512 was a 20B model with a massive:
๐พ 40.9GB BF16 image transformer
But the new Qwen-Image-2.1 comes with โฆ
๐ง 7B visual generator
๐พ 14.2GB BF16
๐ฅ 7.26GB INT8 already available for ComfyUI ๐๐ (Day-1)
This little thing can โฆ
๐จ generate AND edit images
๐ผ๏ธ generate native 2K
๐ซฅ create real transparent RGBA images
๐ฅ use up to 10 reference images
โ๏ธ preserve people/products while editing
๐ค render text
๐ฏ do masked/local edits
Qwen basically took the huge local image model and reduced with a catch โฆ
โ ๏ธ The tiny 7.26GB image model comes with an entire pipeline.
It uses a Qwen3-VL 8B encoder + VAE.
ComfyUI already has a 6.31GB W4A8 encoder, and Qwen supports CPU offloading for smaller GPUs.
Alsoโฆ
๐ซ Qwen-Image-2.1 is NON-COMMERCIAL under the new Qwen Research License.
Thatโs especially notable because Qwen-Image-2512 was Apache 2.0.
Show more
On the SAME 48GB M5 Pro running ๐ง Qwen3.6-35B-A3B @ 210 tps and @ 143 tok/s at 32K context, the multi-agent handling is so much better than mere tokens.
With 32K of context already cached, Splash produced the first token in โก 123 ms vs ๐ 816 ms for oMLX
And when Inco tested 16 simultaneous 32K requests with Qwen3.8-27Bโฆ
Splash handled all 16.
Itโs built to handle heavy agent workloads with lots of context and repeated prompts even with multiple agents at once
w/fast reuse of cached context
Show more
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine!
It's a new open-source inference engine called Splash.
Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model ๐
โ๏ธ model-specific kernels
๐ง hardware-aware memory planning
๐ DFlash2 speculative decoding
๐พ prompt-cache reuse
๐ฅ continuous batching
On the same 48GB M5 Pro running Qwen3.8-27B:
๐ Splash: 74 tok/s
โก oMLX: 38 tok/s
๐ Ollama: 24 tok/s
At 32K context:
๐ Splash: 54 tok/s
And with 4 concurrent requests:
๐ฅ 170 aggregate tok/s
vs 43 tok/s for oMLX
Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max.
And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime.
Requirements
๐ M3 or newer
๐พ 36GB minimum
๐ 48GB+ recommended
Completely different performance because the software stack is optimized around what it is running.
โ ๏ธ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test.
๐
Show more
Check out this part of Athena.
1 Spark can be turned into a local multi-model agent server.
Install both๐ง Qwen3.8 Flash-Next & ๐ง DeepSeek V4 Flash
The Spark can't hold both in its 128GB memory at once, so Athena can checkpoint the current agent/session then ..
๐งน unload one giant model
๐ load the other
โก switch in ~46 sec
๐ keep the same OpenAI/Anthropic API endpoint
Your client just talks to Athena.
Both quantized models scored 91/100 on Athena's tool-use eval:
๐ ๏ธ 69 scenarios
๐งฐ 52 possible tools
๐ chained calls
๐ซ knowing when NOT to call
๐ error recovery
๐ schema-valid JSON
โ ๏ธ Both models had trouble with one tool-output prompt-injection scenario.
Athena is free for personal/research/education use, but the engine itself is currently proprietary/closed-source.
Show more
๐ฃ New Inferenceing Engine Alert!
One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP.
A new engine called Athena recently dropped for Nvidia GB10 systems.
On a single 128GB DGX Spark you can get ...
DeepSeek V4 Flash
โก 8K prefill: 1,126 tps
๐ 256K: 948 tps
๐ Decode
@256K: 19.4 tps
Qwen3.8 Flash-Next
โก 8K prefill: 1,071 tps
๐ 256K: 961 tps
๐ Decode
@256K: 32.1 tps
Going from 8K โ 256K barely rocks on prefill performance.
And Athena caches long conversations to disk. ๐๐
A 141,519-token conversation reportedly restores in:
โก 2.1 seconds vs ~2 min 20 sec to process it fresh.
It also includes
โ
speculative decoding
โ
OpenAI + Anthropic-compatible API
โ
tool calling
โ
Qwen image/document input
โ
persistent agent context
โ
Docker install
โ
switch models without changing clients
This is what I want from DGX Spark.
๐ฏ Huge models + huge context + usable speed on 1 box sitting on your desk. ๐
๐ Link in ALT
Show more
The more I look at setups like this, the more I think OCuLink changes the perspective on Local AI mini-PCs.
That old RTX 2080 Super doesn't need to hold your giant LLM to be useful.
Give the NVIDIA GPU the CUDA-friendly jobs like ๐๏ธ vision models and ๐จ image generation and ...
mini-PC's CPU/iGPU/system RAM handles the larger model or another workload when 8GB isn't enough.
Show more
Loving this OCuLink setup with the GMKtec EVO-T1 paired with an RTX 2080 Super. An Intel mini-PC gets dedicated Nvidia CUDA + 8GB of VRAM over PCIe 4.0 ร4 OCuLink. Local AI possibilities are endless!
Could run separate models/workloads across the iGPU/CPU and NVIDIA GPU or experiment with model offloading, or even play with disaggregated inference!
Show more
๐ฃ New Inferenceing Engine Alert!
One DGX Spark is now running DeepSeek V4 Flash + Qwen3.8 Flash-Next at 262K context with ~1,000 tps PP.
A new engine called Athena recently dropped for Nvidia GB10 systems.
On a single 128GB DGX Spark you can get ...
DeepSeek V4 Flash
โก 8K prefill: 1,126 tps
๐ 256K: 948 tps
๐ Decode
@256K: 19.4 tps
Qwen3.8 Flash-Next
โก 8K prefill: 1,071 tps
๐ 256K: 961 tps
๐ Decode
@256K: 32.1 tps
Going from 8K โ 256K barely rocks on prefill performance.
And Athena caches long conversations to disk. ๐๐
A 141,519-token conversation reportedly restores in:
โก 2.1 seconds vs ~2 min 20 sec to process it fresh.
It also includes
โ
speculative decoding
โ
OpenAI + Anthropic-compatible API
โ
tool calling
โ
Qwen image/document input
โ
persistent agent context
โ
Docker install
โ
switch models without changing clients
This is what I want from DGX Spark.
๐ฏ Huge models + huge context + usable speed on 1 box sitting on your desk. ๐
๐ Link in ALT
Show more
Interesting HuggingFace post.
What if the easiest way to make long-context Local AI faster is to simply STOP feeding it everything?
A new HF forum experiment tested 600 distractor-heavy long-context prompts with Qwen2.5-7B 4-bit.
Instead of sending the entire ~14K-token context, a MiniLM semantic selector kept only the most relevant ~60%.
The result ...
๐ Input tokens
13,124 โ 7,702
โก๏ธ -41%
โก Prefill
871 ms โ 463 ms
โก๏ธ -47%
๐ End-to-end latency
944 ms โ 519 ms
โก๏ธ -45%
๐พ Peak VRAM
11.71 GiB โ 9.16 GiB
And here's the weird part...
๐ฏ Token F1 actually IMPROVED:
0.534 โ 0.581
So the model got:
โ
less context
โ
lower VRAM use
โ
almost 2ร faster prefill
โ
lower total latency
โ
BETTER answers
Why?
Because more context isn't always better context.
Coding agents and RAG systems often drag around enormous histories filled with irrelevant junk.
โ ๏ธ This was a controlled distractor-heavy HotpotQA stress test, not proof that deleting 40% of every real-world context will improve accuracy.
๐ Link in ALT.
Show more
Intel's 32GB Arc Pro B70 finally got a proper Qwen3.8-27B Local AI workout.
So this is why the 32GB VRAM is important, even from an Intel GPU.
Here is Gigazine's setup. Windows 11 using Unsloth Desktop + llama.cpp ...
๐ง Qwen3.8-27B Q4_K_XL
๐พ 30.0GB VRAM + ~1GB shared RAM
๐ 23.4 tps
Then they tried Q8:
๐ง Qwen3.8-27B Q8_K_XL
๐พ 30.2GB VRAM + ~1.5GB shared RAM
โก 15.9 tps
That's a full 27B-class model running at very usable speeds on an Intel GPU.
And also tested text-to-image models.
๐จ Z-Image-Turbo
1024ร1024 / 8 steps
โก๏ธ ~5.3 sec/image after warmup
๐ฅ MiniMax H3 video also ran
โฆbut these are much slower, demonstrating where Nvidia's more mature AI software stack still matters.
The hardware setup...
๐ฎ Arc Pro B70
๐พ 32GB GDDR6
โก 608 GB/s bandwidth
๐ 230W
And here's the interesting value angle.
In Japan, the tested ASRock B70 was selling for:
๐ฐ ยฅ298,054
while many RTX 5090s were selling above:
๐ฐ ยฅ900,000
โ ๏ธ That's Japan-specific street pricing, NOT a universal B70-vs-5090 price comparison.
But THIS is why the B70 interests me for Local AI.
It won't beat a 5090 but it's giving you 32GB of VRAM at a much lower entry point.
23.4 tok/s on Qwen3.8-27B Q4 good enough for you?
๐ GIGAZINE review in ALT
Show more
Okay, so specialized decision models are becoming the rage.
Bespoke Labs just released Nimble, an open 9B decision model built from Qwen3.5-9B.
And apparently they built it in ONE DAY.
The recipe:
๐ง Qwen3.5-9B
๐ฏ LoRA fine-tune
๐ only 2,676 training examples
โ no Jev distillation
โ no RL
โ
open model
โ
open data
โ
open training recipe
Instead of generating explanations or JSON, Nimble scores the allowed answers:
YES / NO โ A / B / C โ route 1 / 2 / 3 โ severity 1โ5
Then returns the choice + probabilities.
Example of Bespoke's 324-example held-out test:
Qwen3.5-9B โ 66.36%
Qwen3.8-27B โ 84.88%
๐ฅ Nimble-9B โ 90.12%
Jev โ 93.21%
And because it isn't generating a bunch of tokens:
โก ~106 ms median on H100
๐ ~444 ms median on an M5 Pro 64GB ๐ ๐
๐ฏ So this is the key, giant reasoning model doesn't need to answer every tiny agent decision.
Let the big model think, then let a small local model handle the ...
๐ ๏ธ tool selection
๐งญ routing
โ
verification
๐จ policy decisions
๐ scoring
๐ retry / stop
Fast. Local. No API call.
โ ๏ธ Bespoke says this is a narrow synthetic 324-example evaluation, NOT a standardized System One benchmark.
๐ GH: /bespokelabsai/nimble
Show more
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine!
It's a new open-source inference engine called Splash.
Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model ๐
โ๏ธ model-specific kernels
๐ง hardware-aware memory planning
๐ DFlash2 speculative decoding
๐พ prompt-cache reuse
๐ฅ continuous batching
On the same 48GB M5 Pro running Qwen3.8-27B:
๐ Splash: 74 tok/s
โก oMLX: 38 tok/s
๐ Ollama: 24 tok/s
At 32K context:
๐ Splash: 54 tok/s
And with 4 concurrent requests:
๐ฅ 170 aggregate tok/s
vs 43 tok/s for oMLX
Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max.
And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime.
Requirements
๐ M3 or newer
๐พ 36GB minimum
๐ 48GB+ recommended
Completely different performance because the software stack is optimized around what it is running.
โ ๏ธ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test.
๐
Show more
Nvidia dropped an official DeepSeek-V4.1-Flash NVFP4 build.
And this thing is BIG.
DeepSeek V4.1 Flash stats
๐ง 552B backbone
๐ +196B Engram conditional memory
โก only 8B active during prefill
๐ 16B active during decode
๐๏ธ native vision
๐ 1 MILLION token context
๐งฉ 384 routed experts across 40 layers
๐ MIT license
Nvidia has converted its routed MoE experts to NVFP4 W4A4 specifically for Blackwell GPUs.
This is not a compressed giant model down to 4-bit.
DeepSeek's experts were already stored in MXFP4, Nvidia instead converts them to its Blackwell-friendly NVFP4 format.
And because NVFP4 uses finer scaling, the checkpoint actually gets slightly larger.
๐พ Source: ~476 GiB
๐พ NVIDIA NVFP4: ~492 GiB
48 safetensor shards. ๐ณ
So why bother? Because NVIDIA is optimizing how those 4-bit experts execute on Blackwell.
And impressively, NVIDIA's evaluations show basically no obvious quality collapse from the conversion.
For example:
๐ง GPQA Diamond
91.04 โ 91.29
๐ป SciCode
54.40 โ 55.84
๐ ๏ธ Terminal-Bench 2.1
81.60 โ 82.16
๐๏ธ MMMU-Pro
74.05 โ 73.70
Some slightly up.
Some slightly down.
Essentially benchmark parity.
And it already has:
โ
vLLM support
โ
SGLang support
โ
reasoning parser
โ
tool calling
โ
image input
โ
1M context
โ ๏ธ Nvidia validated it on 4ร GB300 GPUs so not a local model (yet). The checkpoint is still ~492 GiB.
๐ HF: /nvidia/DeepSeek-V4.1-Flash-NVFP4
Show more
๐คฏ This is like having Jev at home, a AI reflex system running locally.
Someone is turning Google's open DiffusionGemma 26B-A4B into a local version of the new โSystem Oneโ AI idea inside vLLM.
Remember Jev? It is a decision LLM. Instead of asking an LLM to generate paragraphs, System One models are designed to make FAST structured decisions:
๐ฏ yes / no
๐งญ route A / B / C
๐ ๏ธ which tool to call
๐จ severity 1โ4
๐ classify / score / choose
And DiffusionGemma has an important role because a normal autoregressive LLM takes Prompt โ token โ answer
DiffusionGemma is Prompt โ predefined answer slots โ fill multiple decisions in parallel
Because it operates over a token canvas with bidirectional attention instead of being forced to generate everything strictly left-to-right.
The vLLM patch lets it return
โ
bounded choices
โ
probabilities
โ
confidence / uncertainty
โ
multiple decisions simultaneously
And it runs locally.
On ONE DGX Spark the developer reports
โก 1 request: ~120 ms
๐ concurrency 32: ~54 requests/sec
๐ง 3 decisions/request
๐ฅ ~162 decisions/sec
Then the developer found each question can require only around 3 canvas tokens:
๐ข index
โ placeholder
โ๏ธ separator
Meaning as many as ~85 questions could fit into 1 diffusion canvas.
And subsequent optimization reportedly cut decision time another:
โก ~20โ40% with no change in decision quality.
The model stats ...
๐ง 25.2B total parameters
โก 3.8B active
๐พ NVIDIA NVFP4: ~18.9GB
๐ฎ Can fit on a 24GB NVIDIA GPU
๐ Open weights
โ ๏ธ vLLM PR #
57250# is still OPEN ๐ not merged.
And this is Jev-like functionality. It is NOT evidence that DiffusionGemma matches Jev's intelligence or calibration.
๐ /vllm-project/vllm/pull/57250
Show more
Loving this OCuLink setup with the GMKtec EVO-T1 paired with an RTX 2080 Super. An Intel mini-PC gets dedicated Nvidia CUDA + 8GB of VRAM over PCIe 4.0 ร4 OCuLink. Local AI possibilities are endless!
Could run separate models/workloads across the iGPU/CPU and NVIDIA GPU or experiment with model offloading, or even play with disaggregated inference!
Show more
Been using OpenAI's new Image Gen 2.5. Here is one of my favorites generated with the new tool.
๐คฏ Remember back in July I called Ling-3.0-flash one of the more interesting Local AI models?
... because Ling-3.0-flash combined 124B-scale capacity with only ~5B active parameters per token, strong frontierish benchmarks, and a design that looked promising for local agents.
My July post ended with one BIG problem ๐ localmaxxers couldn't download the weights or load it into llama.cpp.โ
Well... that part changed. ๐ฅ
@TheInclusionAI has now released the weights under MIT, an official GGUF exists for Ling-3.0-flash, and they've just released Ling-3.0-flash-VL.
The new VL model adds:
๐ Images
๐ฅ Video
๐ค Visual/computer-use agents
๐ง 124B total / ~5.5B active
๐ up to 1M context
โ๏ธ MIT
A community Ling-3.0-flash-VL Q4_K_M GGUF is already ~79.3GB
~1.7GB vision projector.
So roughly 81GB for a 124B multimodal MoE.
That puts it in range of these devices ...
๐ง 128GB Strix Halo
๐ high-memory Apple Silicon
๐ฎ large/multi-GPU home rigs
โ ๏ธ The VL GGUF is still experimental and currently uses a patched llama.cpp, so this is not yet a clean stock-llama.cpp experience.
But this is exactly the update I wanted back in July, weights on our own machines. ๐ฅ
Now somebody run that ~81GB VL stack on a 128GB Strix Halo and give us the tps. ๐
๐ HF inclusionAI/Ling-3.0-flash-VL
Show more
๐คฏ Remember the weird U.S.-built diffusion LLM I posted about that broke 1,000 tok/s on Nvidia GPUs?
It got a BIG upgrade.
@_inception_ai released Mercury 2.5 today.
๐ฏ Instead of generating one token after another like a normal LLM, Mercury uses diffusion to generate/refine multiple tokens in parallel.
And Inception says the new model delivers the following ๐
๐ 1,107 tok/s ๐
๐ง +40% intelligence vs Mercury 2
๐ 260K context
๐ค Tunable reasoning
๐ง Parallel tool calling
๐ Structured JSON
๐ต $0.20/M input / $0.75/M output ๐
๐ฏInception says this is the largest diffusion language model ever trained.
That ~1,100 tok/s isn't coming from Groq or Cerebras-style custom inference hardware - this is an important point!
It's running on widely available NVIDIA GPU infrastructure.
That's the big architectural bet, traditional LLM ๐
token โ token โ token โ token
Diffusion LLM ๐
many tokens โ refine them together โ answer
โ ๏ธ Caveat, of course! ๐ Mercury 2.5 is closed-weights.
Show more
Today, weโre introducing Mercury 2.5, the most capable diffusion LLM on the market.
It offers a 40% jump in intelligence over Mercury 2, and runs over 1,100 tokens/sec on widely-available
@NVIDIAAI GPUs.
Itโs available today on our API,
@OpenRouter, and
@Baseten.
Contact us to evaluate Mercury 2.5 for production:
Show more