가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

LaurieWired
@lauriewired
researcher @google; serial complexity unpacker; ex @ msft & aerospace
가입 January 2023
303 팔로잉 중    163.2K 팬
If you’re a CS student, I’d suggest writing your thesis around SDCs (silent data corruptions) and efficient algorithmic fault tolerance. We’re quickly going to need to accept the (un)reliability of computing systems again! Every indicator (density, lower voltage gating, raw scale) is moving towards SDCs becoming a more serious issue over the next decade. And let’s not forget when all of this stuff starts to go in space and we have to deal with cosmic events on “not-super-rad-hard” CPUs+GPUs! No, you can’t solve it all in hardware. It’s prohibitively expensive; and the long tail of “potential SDCs” is just too long. One bad apple ruins your pie. One bad GPU can inject repeated corruptions into otherwise noise-tolerant workloads (*cough* AI training). If you continuously accept a corrupted value from say…a bad tensor core, that bad value can quickly spread across the whole cluster. If it happens again and again and again and no one notices; it’s like putting a little bit of spin on a bowling ball. Eventually, the overall trajectory ends up widely different! When you start to imagine this stuff being in space (cosmic bit flips), and quantized(!), each remaining bit carries more critical information. Long term we’re going to have to accept clever, low-overhead software algorithms. Accepting say, a ~3% perf loss can be *absolutely* worth it if it increases your odds of detecting a mercurial core enough!
더 보기
0
111
6.5K
406
커뮤니티로 전달