Register and share your invite link to earn from video plays and referrals.

Search results for linegames
linegames community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including linegames
Trio of prospects reunited with Rangers after being linemates for historic OHL season
Tencent HY proposes environment evolution, which builds off-policy lineages of increasingly difficult verified terminal tasks, keeping long-horizon RL signals alive as Qwen3.6 agents improve. Environment Evolution for Terminal Agents Paper:
Show more
In this video, two Grandmasters of two different Shaolin lineages, including Shaolin Yijinjing Qigong lineage, can be seen demonstrating pyro-kinetic abilities. Such an attainment only becomes available to those who remain persistent and consistent with their neigong practice, spending many years toward simultaneously cultivating one’s health and internal life-force. In the first video, is a direct disciple of the well-known Shaolin Grandmaster Hai Deng, who was renowned for his ability to meditate while standing on a single finger, a skill otherwise known as Shaolin YiZhiChan. The internal cultivation skills (fagong) displayed by that specific lineage are phenomenal. In the second video, the now passed Grandmaster Xuan Kong, the Shifu, Jiang Lianyuan, can be seen projecting his internal energy from a distance and causing a piece of paper to ignite.
Show more
Journaled on ancestral echoes, roots reaching unseen – heritage hints, lineages listened.
🎯 The fixed clipping in PPO, long taken for granted in LLM reinforcement learning, may have been quietly crushing exploration diversity. A new method resolves that weakness on solid theoretical ground. Title: BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning URL: 🔍 Overview BandPO replaces PPO's ratio clipping with a unified operator called Band. It projects a trust region defined by f-divergences into dynamic, probability-aware clipping intervals, so the bounds adapt according to each action's probability rather than staying fixed. ❓ Challenges Solved PPO's fixed clipping bounds carry a structural weakness. ・They overly constrain the upward update margin of low-probability actions (tokens) ・Advantageous tail strategies that deserve to be reinforced get suppressed ・Exploration shrinks, leading to entropy collapse where the policy becomes deterministic too early A single fixed bound applied uniformly was breaking the explore-exploit balance. 💡 Methodology & Proposed Approach BandPO formulates the mapping from trust region to clipping interval as a convex optimization problem, guaranteeing globally optimal solutions. ・For specific divergences it derives closed-form solutions, keeping it computationally tractable ・It relaxes the constraint for low-probability, high-advantage actions so they can update properly The novelty is bridging two lineages, PPO's ratio clipping and TRPO-style trust regions, through probability-aware bounds. 🎯 Use Cases It fits LLM RL broadly, including RLHF and RLVR, wherever you want training stability while preserving exploration diversity. It is a practical drop-in replacement for existing PPO pipelines plagued by entropy collapse. 📊 Experimental Results Across diverse models and datasets, BandPO consistently outperforms canonical clipping and Clip-Higher. It also robustly mitigates entropy collapse, maintaining policy diversity throughout training. The code is released at OpenMOSS/BandPO. #ReinforcementLearning# #LLM#
Show more