Register and share your invite link to earn from video plays and referrals.

AJ
@ItsmeAjayKV
Bullish on local AI, llm finetuning, abliteration, llama.cpp OSS contributions GPUmaxxing: 1x 3060, 1x 3090 N4 (日本語勉強中) 🎌
683 Following    3.5K Followers
After getting GLM-5.3 flash, we are getting the main model itself. GLM-5.3 Less than 15 hrs to go August has been a crazy month for Local AI LFGGGG🥳
llama.cpp PR #27742# updated q8_0 KV was broken on Qwen3.8-Flash-Next when i tried for first time. Now it's fixed, so take a pull and rebuild. I was running f16 KV. Dropping to q8_0 should let me keep more experts on GPU (less CPU offload) and get more speed.
Show more
Here is the first speed test of @Alibaba_Qwen Qwen3.8-Flash-Next, running on my 3090 😍 @UnslothAI UD-IQ3_XXS: 82GB Total memory: 88GB 24 GB VRAM + 64 GB RAM llama.cpp (pr #27742#) - Total usable context at f16 kv: 158K - Total VRAM usage 23GB Tested at depth up to 128k context. - Avg decode speed: 21 t/s - Peak decode speed: 33 t/s - Peak prefill speed: 168 t/s @ 32k All results were run without MTP or DFLASH
Show more
So we now have 3 flash models DSV4-flash-0731 Qwen3.8-flash-next GLM-5.3-flash I think this is solid, every model should have a faster flash version, this is exactly what i asked for a while ago. So between these three, what do you say is the best when running locally?
Show more
I'm so happy with decode speed of Qwen3.8-Flash-Next on my 3090! It's really usable, cannot wait for next two weeks of optimizations which will push it to 30 - 40t/s. Prefill hurts thoo... 😭
Show more
I woke up and thankfully there is no other release. Back to running Qwen3.8-flash-next
Qwen3.8-Flash-Next on a single RTX PRO 6000 👀 Quant: Unsloth AI UD-Q4_K_XL Model size: 111 GB GPU: RTX PRO 6000, 96 GB GDDR7 Fully on GPU with full context. Benchmark context tested up to ~253K Results: - 40.3 t/s average decode across 8K–253K - 63 t/s peak decode @ 8K - 17.6 t/s decode @ ~253K - 1,657 t/s peak prefill @ 8K All results were run without MTP or DFLASH.
Show more
Before i go to sleep, i'm giving Qwen3.8-flash-next-UD-Q4_K_XL an overnight task via DeepSeek Harness. Will be interesting how it performs, it's certainly not a one-shot type model, but getting really good and detailed results, better than 27b already. BTW this is running on rented pro 6000, let's see!
Show more
Here is the first speed test of @Alibaba_Qwen Qwen3.8-Flash-Next, running on my 3090 😍 @UnslothAI UD-IQ3_XXS: 82GB Total memory: 88GB 24 GB VRAM + 64 GB RAM llama.cpp (pr #27742#) - Total usable context at f16 kv: 158K - Total VRAM usage 23GB Tested at depth up to 128k context. - Avg decode speed: 21 t/s - Peak decode speed: 33 t/s - Peak prefill speed: 168 t/s @ 32k All results were run without MTP or DFLASH
Show more
Here's the pr to watch for Qwen3.8-Next-Flash llama.cpp support. It's from UnslothAI team themselves.
Xiaomi just showed its AI Cube Prototype and this could become a serious GB10 competitor from China 👀 - 3 custom chips: Xring O3, O100, D100 - 200 TOPS NPU - 1.22 TB/s AI memory bandwidth - Up to 160GB unified memory - 150W sustained power - 120B models running locally Xring O100: 1.22TB/s + 330 t/s on a 150w AI box is 🔥 Once it hit's the marked, going to sell like hot cakes.
Show more
0
192
4K
336
Forward to community
Here's a comparison i never thought i would make Qwen-3.8-27b vs Gemini-3.7-flash Yes, you read it right! I don't know what to say, this is soo damn impressive. the difference is hugeee🤯 Take a look at this result 😍 (watch it in highest quality) Same prompt, both one shot on web-chat ui Qwen-3.8-27b produces consistent beautiful, visually stunning results. This is one of the best water surface simulations i have seen across every local models i have run and at the same time beating Gemini-3.7-flash result. And this is running on my 3090. And yes its Q5, not even f8.
Show more
if you're on llama.cpp run Qwen3.8-27b with -spec-default --spec-type draft-mtp Since it does take its sweet time to think, might as well let it think fast. Run with mtp, you don't need a sep drafter, getting 2x speed on decode now, totally worth it. full command i'm using on my 3090 ./build/bin/llama-server -m "/qwen-3.8/Qwen3.8-27B-Q4_K_M.gguf" --host 127.0.0.1 --port 8080 -ngl 999 -fa on --jinja -np 1 -t 12 --alias qwen3.8-27b-q4 --spec-default --spec-type draft-mtp --cache-type-k q8_0 --cache-type-v q8_0
Show more