가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

wh
@nrehiew_
eng primarily, ml mostly, research previously
가입 October 2023
104 팔로잉 중    18.5K 팬
RL Infra time. - nice dispatch strategy that gets rid of long tail stalls - router replay from previous checkpoints - this ends up causing shorter completions to impact early stages of training and off-policy. Their solution is to cap at the dataset level (so datasets with general shorter responses dont influence too much) and a discard scheme - For offpolicyness, they bound the off-policy ratio and loss masking. so pretty standard here - When a new checkpoint is updated, KVs and routers are persistent instead of being recomputed At the final stage they do full vocab OPD on >40 teacher models
더 보기