登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

wh
@nrehiew_
eng primarily, ml mostly, research previously
参加 October 2023
104 フォロー中    18.5K ファン
RL Infra time. - nice dispatch strategy that gets rid of long tail stalls - router replay from previous checkpoints - this ends up causing shorter completions to impact early stages of training and off-policy. Their solution is to cap at the dataset level (so datasets with general shorter responses dont influence too much) and a discard scheme - For offpolicyness, they bound the off-policy ratio and loss masking. so pretty standard here - When a new checkpoint is updated, KVs and routers are persistent instead of being recomputed At the final stage they do full vocab OPD on >40 teacher models
もっと見る