KernelBench update:
i went through a bunch of traces looking for models attempting to use ncu and getting blocked due to the compute provider disabling performance counter permissions (NVreg_RestrictProfilingToAdminUsers=1). this affected rtx pro 6000 (using a compute provider as my own compute was saturated at the time) h100 and b200 runs. im thinking of doing a bunch of reruns but this would cost me an arm and a leg.
im very impressed with models ability to optimize kernels without all the information from ncu profiles. i also wonder if kernel-based rl from labs has ran into this issue and solved it before training, or if models just have this strange intuition about what patterns tend to perform well.
before i do anything, i want to get some feedback from the community on this.