Ngram module (figure from DeepSeek's Engram)
They first ablate at which layer to put it but couldnt find a clear winner. As such they just put one at layer 2 which allows it to be prefetched from CPU together with layer 1 computation
From the figures in this section, this seems to be the biggest area where loss is not a good signal on its own. Downstream loss goes down as the vocab size increases, but same isnt true for evals