๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

AJ
@ItsmeAjayKV
Bullish on local AI, llm finetuning, abliteration, llama.cpp OSS contributions GPUmaxxing: 1x 3060, 1x 3090 N4 (ๆ—ฅๆœฌ่ชžๅ‹‰ๅผทไธญ) ๐ŸŽŒ
๊ฐ€์ž… September 2016
683 ํŒ”๋กœ์ž‰ ์ค‘    3.5K ํŒฌ
Qwen3.8-Flash-Next on a single RTX PRO 6000 ๐Ÿ‘€ Quant: Unsloth AI UD-Q4_K_XL Model size: 111 GB GPU: RTX PRO 6000, 96 GB GDDR7 Fully on GPU with full context. Benchmark context tested up to ~253K Results: - 40.3 t/s average decode across 8Kโ€“253K - 63 t/s peak decode @ 8K - 17.6 t/s decode @ ~253K - 1,657 t/s peak prefill @ 8K All results were run without MTP or DFLASH.
๋” ๋ณด๊ธฐ