Register and share your invite link to earn from video plays and referrals.

Search results for RLTraining
RLTraining community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including RLTraining
Running technique retraining can prevent surgery for chronic exertional compartment syndrome (CECS). 10 runners scheduled for fasciotomy retrained. After 6 weeks, muscle pressure ↓ ~50%, pain eased, and all returned to running without surgery.
Show more
🔎 A paper that pinpoints the hidden reason RL training for search agents stalls partway. Title: Harness-G: A Graph-Structured Harness for Search Agents URL: ❓ Why does training collapse? 💡 It's "retrieval-equivalence collapse." The policy keeps generating differently-worded queries that fetch the same evidence, so same evidence → same answer → same reward, within-group advantages vanish, and the training signal dries up. ❓ How does Harness-G fix it? 💡 It stops free-form query generation and turns it into menu selection over a paragraph-sentence-entity graph built from the corpus. The policy picks action IDs, not strings. Being finite, verifiable, and previewable, it preserves diversity in what actually gets retrieved. ❓ What is the credit assignment (SNC)? 💡 A frozen answerer previews how much an action raises the gold-answer probability, scored against alternatives (frontier-relative). Non-myopic payoffs like "find the bridge entity first" propagate back through provenance edges (enablement). No extra rollouts needed. ❓ Does it work? 💡 Across six QA benchmarks it beats Graph-R1 by +10.74 at 1.5B and +3.98 at 3B, best at both scales, especially on multi-hop, with $0 API cost to build the graph. #SearchAgents# #ReinforcementLearning#
Show more
We're hiring at Gradient. Building open-source environment infrastructure for our distributed RL training stack — reproducible, scalable to thousand-GPU runs Looking for 1–2 RL Environments engineers / tech leads: You've designed verifiers, built sandboxes for agentic RL rollouts, or shipped RL training data pipelines that survived contact with real training. Domain depth in math, code, agent, tool, or GUI is a plus. PhD not required. Also hiring research interns: PhD / Masters students with hands-on RLHF / RLVR / GRPO / DPO / agentic RL experience. Open-source footprint matters more than paper count. Most intern roles convert post-grad. No age cap. Founding-team-level equity for the right people. DMs open.
Show more
We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, with fast snapshot, resume, and fork support for large-scale parallel agent workflows. Explore on GitHub:
Show more
0
57
3.4K
444
Forward to community
🤗 LeLab now keeps itself up to date! 🦾 We shipped an in-app popup that alerts you the moment a newer version is on GitHub - plus same upgrades based on users feedback: - Import any external/Hub model to run & rollout, no retraining - Per-camera codec & backend control when recording - Warning when your GPU is present but CUDA isn't being used - Real camera names on Windows & Linux - Smoother calibration & teleop (cameras stay off until you turn them on 🛠️ Update now with: uv tool install --force git+ GitHub: Docs:
Show more
Can someone explain the intuition behind SAEs steering the activations instead of weights? Why would you steer the result instead of the controls? Is it purely to avoid the (arguably gigantic and hard to repro) messiness of retraining?
Show more
Congrats to the @cursor_ai team on the launch of Composer 2! We are proud to see Kimi-k2.5 provide the foundation. Seeing our model integrated effectively through Cursor's continued pretraining & high-compute RL training is the open model ecosystem we love to support. Note: Cursor accesses Kimi-k2.5 via @FireworksAI_HQ ' hosted RL and inference platform as part of an authorized commercial partnership.
Show more
0
517
20.4K
1.4K
Forward to community