Register and share your invite link to earn from video plays and referrals.

Jian Hu
@hijkzzz
I'm a RLer + MLSyser / 2 + NLPer / 2.
7 Following    633 Followers
FlashREINFORCE: to our knowledge, the first open-source critic-free, single-rollout async LLM RL with 6,000+ stable updates. One-Batch REINFORCE + Sequence Trust Region + Sample-Mean Optimization. Paper: Code:
Show more
1/ Still looking for a minimalist, high-performance framework for agentic RL research? Meet Molt — an agentic-first, PyTorch-native reinforcement learning framework with roughly 9K lines of RL code for 700B models. ⭐
Show more