Register and share your invite link to earn from video plays and referrals.

wh
@nrehiew_
eng primarily, ml mostly, research previously
Joined October 2023
104 Following    18.5K Followers
RL Infra time. - nice dispatch strategy that gets rid of long tail stalls - router replay from previous checkpoints - this ends up causing shorter completions to impact early stages of training and off-policy. Their solution is to cap at the dataset level (so datasets with general shorter responses dont influence too much) and a discard scheme - For offpolicyness, they bound the off-policy ratio and loss masking. so pretty standard here - When a new checkpoint is updated, KVs and routers are persistent instead of being recomputed At the final stage they do full vocab OPD on >40 teacher models
Show more