๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
๊ฐ€์ž… May 2018
2.1K ํŒ”๋กœ์ž‰ ์ค‘    11.7K ํŒฌ
New free learning path on the Red Hat Developer Sandbox: compress, serve, and benchmark a model with @vllm_project, hands-on in Jupyter, no GPU needed. Here's what you'll actually do: โšก Quantize Qwen3 to W4A16 with LLM Compressor using GPTQ. Measure the result: 42% smaller, 8.2% perplexity increase. Learn to decide if that tradeoff fits your use case. ๐Ÿš€ Connect to a running vLLM server and send requests via the OpenAI-compatible API. Watch 5 concurrent requests handled in real time. See prefix cache queries increment live via the Prometheus metrics endpoint. ๐Ÿ“Š Run a GuideLLM benchmark: TTFT, inter-token latency, and E2E latency at p50, p95, and p99. Run Hellaswag with lm_eval. Cross-reference with the published model card to make a deployment decision backed by numbers. Less than an hour to complete. Free account. Built by @cedricclyburn and Michael Santos. ๐Ÿ™
๋” ๋ณด๊ธฐ