가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Alexey Fateev
@superalesha
⚡I benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. ❤️ - 2xDGX Spark 🚀96GB VRAM | Local AI
가입 January 2026
325 팔로잉 중    3.7K
I rented an RTX PRO 6000 Blackwell (96 GB) for 10 hours to answer one question: what is better for local agents: vLLM or llama.cpp. The same Qwen 3.6 (27 billion parameters), closest 4-bit quants, byte-identical dialogues of 12 turns, 26 measured slices. llama.cpp was serving in 83 seconds. vLLM took 609 but in everything else, vLLM won almost across the board, haha
더 보기