注册并分享邀请链接,可获得视频播放与邀请奖励。

Tencent Hy
@TencentHunyuan
Tencent's foundation model for text, image, video, and 3D generation.
加入 July 2024
8 正在关注    53K 粉丝
⚡️ As LLM reinforcement learning scales to larger GPU clusters and more training data, training efficiency becomes a first-order concern. Our new research revisits classical critical-batch-size theory and extends it to online LLM RL, where the model generates its own training data and rollout generation and training scale differently. Across GRPO and PPO, we find that learning-rate retuning can preserve learning per response over a bounded range of batch sizes. On fixed hardware, scaling up the batch size improves PPO generation-stage throughput by up to 2.29×, while our best measured GRPO configuration reaches the same validation target in 29% less time. 🚀 Read the full research:
显示更多
0
10
192
8
转发到社区