註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

AJ
@ItsmeAjayKV
Bullish on local AI, llm finetuning, abliteration, llama.cpp OSS contributions GPUmaxxing: 1x 3060, 1x 3090 N4 (日本語勉強中) 🎌
加入 September 2016
683 正在關注    3.5K 粉絲
Qwen3.8-Flash-Next on a single RTX PRO 6000 👀 Quant: Unsloth AI UD-Q4_K_XL Model size: 111 GB GPU: RTX PRO 6000, 96 GB GDDR7 Fully on GPU with full context. Benchmark context tested up to ~253K Results: - 40.3 t/s average decode across 8K–253K - 63 t/s peak decode @ 8K - 17.6 t/s decode @ ~253K - 1,657 t/s peak prefill @ 8K All results were run without MTP or DFLASH.
顯示更多