登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

wh
@nrehiew_
eng primarily, ml mostly, research previously
参加 October 2023
104 フォロー中    18.5K ファン
Sparse Attention with Indexer also seems rather standard. Essentially, we use an indexer at block level in a compressed latent space. For training, the indexer is asked to predict the full attention score 1) Distill the dense attention scores into the indexer 2) KL loss on the indexer vs the full attention teacher
もっと見る