It’s fairly well known frontier labs use value functions now, but we still don't have good open value function infra. Here's something cool from my ongoing internship
@MistralAI with
@laurence_ai getting value functions to work
Introducing - Prime Values
> Clean hackable, independent abstractions built on top of prime-rl by
@PrimeIntellect
> First-class asynchronous value trainer and evaluator nodes, with no trainer bottleneck
> Native value warmup support
> Streaming replay buffer feeds value model exploiting its greater staleness/reuse tolerance
Defaults validated to match or outperform mean-baseline GRPO on both single-turn and multi-turn tasks