註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

NVIDIA AI
@NVIDIAAI
Teaching your AI new tricks.
加入 June 2016
898 正在關注    343.9K 粉絲
A lot of work goes into serving a model efficiently. For the Nemotron 3 Ultra NIM, our engineers tuned caching, memory, parallelism, decoding and more. On four B200 GPUs, those optimizations supported up to 2.5x more concurrent users while maintaining 50 TPS/user. Read the engineering deep dive →
顯示更多
0
31
344
39
轉發到社區