註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Elliot Arledge
@elliotarledge
@infinity_ai_ | made the 12 hr CUDA course | | @shipfr8 alum
加入 November 2022
352 正在關注    43.9K 粉絲
KernelBench update: i went through a bunch of traces looking for models attempting to use ncu and getting blocked due to the compute provider disabling performance counter permissions (NVreg_RestrictProfilingToAdminUsers=1). this affected rtx pro 6000 (using a compute provider as my own compute was saturated at the time) h100 and b200 runs. im thinking of doing a bunch of reruns but this would cost me an arm and a leg. im very impressed with models ability to optimize kernels without all the information from ncu profiles. i also wonder if kernel-based rl from labs has ran into this issue and solved it before training, or if models just have this strange intuition about what patterns tend to perform well. before i do anything, i want to get some feedback from the community on this.
顯示更多