註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

wh
@nrehiew_
eng primarily, ml mostly, research previously
加入 October 2023
104 正在關注    18.5K 粉絲
RL Infra time. - nice dispatch strategy that gets rid of long tail stalls - router replay from previous checkpoints - this ends up causing shorter completions to impact early stages of training and off-policy. Their solution is to cap at the dataset level (so datasets with general shorter responses dont influence too much) and a discard scheme - For offpolicyness, they bound the off-policy ratio and loss masking. so pretty standard here - When a new checkpoint is updated, KVs and routers are persistent instead of being recomputed At the final stage they do full vocab OPD on >40 teacher models
顯示更多