Register and share your invite link to earn from video plays and referrals.

AJ
@ItsmeAjayKV
Bullish on local AI, llm finetuning, abliteration, llama.cpp OSS contributions GPUmaxxing: 1x 3060, 1x 3090 N4 (日本語勉強中) 🎌
Joined September 2016
683 Following    3.5K Followers
llama.cpp PR #27742# updated q8_0 KV was broken on Qwen3.8-Flash-Next when i tried for first time. Now it's fixed, so take a pull and rebuild. I was running f16 KV. Dropping to q8_0 should let me keep more experts on GPU (less CPU offload) and get more speed.
Show more
Here is the first speed test of @Alibaba_Qwen Qwen3.8-Flash-Next, running on my 3090 😍 @UnslothAI UD-IQ3_XXS: 82GB Total memory: 88GB 24 GB VRAM + 64 GB RAM llama.cpp (pr #27742#) - Total usable context at f16 kv: 158K - Total VRAM usage 23GB Tested at depth up to 128k context. - Avg decode speed: 21 t/s - Peak decode speed: 33 t/s - Peak prefill speed: 168 t/s @ 32k All results were run without MTP or DFLASH
Show more