註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

DailyPapers
@HuggingPapers
Tweeting interesting papers submitted at Submit your own at and link models/datasets/demos to it!
加入 March 2025
4 正在關注    21.6K 粉絲
FlowBalance: verifier-grounded self-improvement for reasoning models Improves math reasoning by +2.12 avg over GRPO on Qwen3-8B, with faster training, better stability, and higher solution diversity.
顯示更多