Register and share your invite link to earn from video plays and referrals.

Siddarth
@siddarthv66
Post-training @MistralAI, PhD @Mila_Quebec | RLing LLMs
850 Following    2.3K Followers
It’s fairly well known frontier labs use value functions now, but we still don't have good open value function infra. Here's something cool from my ongoing internship @MistralAI with @laurence_ai getting value functions to work Introducing - Prime Values > Clean hackable, independent abstractions built on top of prime-rl by @PrimeIntellect > First-class asynchronous value trainer and evaluator nodes, with no trainer bottleneck > Native value warmup support > Streaming replay buffer feeds value model exploiting its greater staleness/reuse tolerance Defaults validated to match or outperform mean-baseline GRPO on both single-turn and multi-turn tasks
Show more