가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

wh
@nrehiew_
eng primarily, ml mostly, research previously
가입 October 2023
104 팔로잉 중    18.5K 팬
Ngram module (figure from DeepSeek's Engram) They first ablate at which layer to put it but couldnt find a clear winner. As such they just put one at layer 2 which allows it to be prefetched from CPU together with layer 1 computation From the figures in this section, this seems to be the biggest area where loss is not a good signal on its own. Downstream loss goes down as the vocab size increases, but same isnt true for evals
더 보기