I rented an RTX PRO 6000 Blackwell (96 GB) for 10 hours to answer one question: what is better for local agents: vLLM or llama.cpp. The same Qwen 3.6 (27 billion parameters), closest 4-bit quants, byte-identical dialogues of 12 turns, 26 measured slices.
llama.cpp was serving in 83 seconds.
vLLM took 609
but in everything else, vLLM won almost across the board, haha
顯示更多