가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

AJ
@ItsmeAjayKV
Bullish on local AI, llm finetuning, abliteration, llama.cpp OSS contributions GPUmaxxing: 1x 3060, 1x 3090 N4 (日本語勉強中) 🎌
가입 September 2016
683 팔로잉 중    3.5K
llama.cpp PR #27742# updated q8_0 KV was broken on Qwen3.8-Flash-Next when i tried for first time. Now it's fixed, so take a pull and rebuild. I was running f16 KV. Dropping to q8_0 should let me keep more experts on GPU (less CPU offload) and get more speed.
더 보기
Here is the first speed test of @Alibaba_Qwen Qwen3.8-Flash-Next, running on my 3090 😍 @UnslothAI UD-IQ3_XXS: 82GB Total memory: 88GB 24 GB VRAM + 64 GB RAM llama.cpp (pr #27742#) - Total usable context at f16 kv: 158K - Total VRAM usage 23GB Tested at depth up to 128k context. - Avg decode speed: 21 t/s - Peak decode speed: 33 t/s - Peak prefill speed: 168 t/s @ 32k All results were run without MTP or DFLASH
더 보기