가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

wh
@nrehiew_
eng primarily, ml mostly, research previously
가입 October 2023
104 팔로잉 중    18.5K 팬
Sparse Attention with Indexer also seems rather standard. Essentially, we use an indexer at block level in a compressed latent space. For training, the indexer is asked to predict the full attention score 1) Distill the dense attention scores into the indexer 2) KL loss on the indexer vs the full attention teacher
더 보기