가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Dylan Patel
@dylan522p
SemiAnalysis AI Infrastructure Research and Consulting DMs are open for consulting, quotes, or to talk shop, Opinions my own
가입 April 2018
1K 팔로잉 중    163.8K 팬
Today we are launching InferenceMAX! We have support from Nvidia, AMD, OpenAI, Microsoft, Pytorch, SGLang, vLLM, Oracle, CoreWeave, TogetherAI, Nebius, Crusoe, HPE, SuperMicro, Dell It runs every day on the latest software (vLLM, SGLang, etc) across hundreds of GPUs, $10Ms of infrastructure is purring every day to create real world LLM Inference benchmarks InferenceMAX answers the major questions of our times with AI Infrastructure. How many Tokens are generated per MW of capacity on different infrastructure? How much does a million tokes cost? What is the real latency vs throughput tradeoff? We have coverage of over 80% of deployed FLOPS globally by covering H100, H200, B200, GB200, MI300X, MI325X, and MI355X. Soon we will be over 99% with Google TPUs and Amazon Trainium being added.
더 보기
0
107
1.7K
138
커뮤니티로 전달