註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ananth Veluvali
@AnanthVeluvali
ars longa, vita brevis | prev. exited founder, @stanford
加入 April 2022
299 正在關注    1.1K 粉絲
I connected devin to some @modal H100s and then left it alone to run its own experiment loop against the GPUs. it found a way to reduce peak memory by up to 46% and latency by up to 52% on a popular OS training-kernel repo. since the triton kernel was already near peak HBM bandwidth, it focused on other things, removing a 2GB allocation that was happening on every chunk of the backward pass. then it ran the benchmarks to confirm the speedup 🤯
顯示更多