๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

AJ
@ItsmeAjayKV
Bullish on local AI, llm finetuning, abliteration, llama.cpp OSS contributions GPUmaxxing: 1x 3060, 1x 3090 N4 (ๆ—ฅๆœฌ่ชžๅ‹‰ๅผทไธญ) ๐ŸŽŒ
๊ฐ€์ž… September 2016
683 ํŒ”๋กœ์ž‰ ์ค‘    3.5K ํŒฌ
Here is the first speed test of @Alibaba_Qwen Qwen3.8-Flash-Next, running on my 3090 ๐Ÿ˜ @UnslothAI UD-IQ3_XXS: 82GB Total memory: 88GB 24 GB VRAM + 64 GB RAM llama.cpp (pr #27742#) - Total usable context at f16 kv: 158K - Total VRAM usage 23GB Tested at depth up to 128k context. - Avg decode speed: 21 t/s - Peak decode speed: 33 t/s - Peak prefill speed: 168 t/s @ 32k All results were run without MTP or DFLASH
๋” ๋ณด๊ธฐ