가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Siddarth
@siddarthv66
Post-training @MistralAI, PhD @Mila_Quebec | RLing LLMs
가입 September 2023
850 팔로잉 중    2.3K
It’s fairly well known frontier labs use value functions now, but we still don't have good open value function infra. Here's something cool from my ongoing internship @MistralAI with @laurence_ai getting value functions to work Introducing - Prime Values > Clean hackable, independent abstractions built on top of prime-rl by @PrimeIntellect > First-class asynchronous value trainer and evaluator nodes, with no trainer bottleneck > Native value warmup support > Streaming replay buffer feeds value model exploiting its greater staleness/reuse tolerance Defaults validated to match or outperform mean-baseline GRPO on both single-turn and multi-turn tasks
더 보기