注册并分享邀请链接,可获得视频播放与邀请奖励。

AJ
@ItsmeAjayKV
Bullish on local AI, llm finetuning, abliteration, llama.cpp OSS contributions GPUmaxxing: 1x 3060, 1x 3090 N4 (日本語勉強中) 🎌
加入 September 2016
683 正在关注    3.5K 粉丝
Here is the first speed test of @Alibaba_Qwen Qwen3.8-Flash-Next, running on my 3090 😍 @UnslothAI UD-IQ3_XXS: 82GB Total memory: 88GB 24 GB VRAM + 64 GB RAM llama.cpp (pr #27742#) - Total usable context at f16 kv: 158K - Total VRAM usage 23GB Tested at depth up to 128k context. - Avg decode speed: 21 t/s - Peak decode speed: 33 t/s - Peak prefill speed: 168 t/s @ 32k All results were run without MTP or DFLASH
显示更多
0
23
181
5
转发到社区