註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Siddarth
@siddarthv66
Post-training @MistralAI, PhD @Mila_Quebec | RLing LLMs
加入 September 2023
850 正在關注    2.3K 粉絲
It’s fairly well known frontier labs use value functions now, but we still don't have good open value function infra. Here's something cool from my ongoing internship @MistralAI with @laurence_ai getting value functions to work Introducing - Prime Values > Clean hackable, independent abstractions built on top of prime-rl by @PrimeIntellect > First-class asynchronous value trainer and evaluator nodes, with no trainer bottleneck > Native value warmup support > Streaming replay buffer feeds value model exploiting its greater staleness/reuse tolerance Defaults validated to match or outperform mean-baseline GRPO on both single-turn and multi-turn tasks
顯示更多
0
10
247
35
轉發到社區