登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

wh
@nrehiew_
eng primarily, ml mostly, research previously
参加 October 2023
104 フォロー中    18.5K ファン
Ngram module (figure from DeepSeek's Engram) They first ablate at which layer to put it but couldnt find a clear winner. As such they just put one at layer 2 which allows it to be prefetched from CPU together with layer 1 computation From the figures in this section, this seems to be the biggest area where loss is not a good signal on its own. Downstream loss goes down as the vocab size increases, but same isnt true for evals
もっと見る