When we started
@cerebras, we knew we would have to solve difficult problems in yield, power, and cooling. What became clear was that making wafer-scale work would require rethinking almost the entire system.
Today, we answered questions from the community about those engineering decisions, how we run larger models, and what it takes to program the hardware.
2:15 Why hasn't NVIDIA gone wafer-scale?
4:10 How we build a working chip out of an imperfect wafer
8:07 What happens when a model won't fit on one wafer?
11:34 Why we chose SRAM over HBM
14:26 How much of a barrier is NVIDIA's CUDA moat?