Register and share your invite link to earn from video plays and referrals.

Search results for RLツアー
RLツアー community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including RLツアー
RL teaches models to work longer, but reasoning is dependent on domain-specific post-training. Baseten's Head of Model Training @oneill_c sat down with @dwarkesh_sp to explain horizon generalization and what's next at the frontier. Full episode here:
Show more
"RL will kill us" "no it won't. have you trained a model?" "no. have you?" "no"
RL is the way to go for coding, math, etc, it’s now pretty clear, but what about writing/prose? What is the solution to really improve models there?
RL hold-out-1 experiment im curious about to study how model specialization relates to model capacity and data allocation for how many domains to specialize for a common way to get a model to specialize in N domains is Multi-Teacher-On-Policy Distillation train N specialized teachers, do OPD with routing for a set of prompts to teach a model all those abilities for a given domain like data viz -> if we remove X% of the teachers (ie. Don’t specialize on those skills), does our performance on data viz increase? how is this affected by which domains get left out? does it suffer if similar domains are removed but benefit if very different domains are? we still see vertical focused specialized models like GPT-Cyber clearly looks like allocating a lot of data and parameters for a given vertical boosts perf in that vertical more than a general purpose mixture points towards a future where the specialist models always win for high value domains
Show more
RL reward is a very lossy way to codify real-world feedback. Thus, it’s very important for ML engineers to actually use the products daily.
RL fine-tuning is now live for @nvidiaai Nemotron 3 on Fireworks, starting with Nemotron 3 Super (LoRA). Train with GRPO and serve the model in one place. We price by GPU-hour, not per token, so long multi-turn rollouts don't blow up your bill. Training shapes →
Show more
Anthropic RL-trained Claude Haiku 4.5 to be an alignment auditor inside Petri, their production auditing scaffold. Training Alignment Auditors via Reinforcement Learning Paper:
Show more
In RL training, a vLLM rollout engine and a Megatron trainer can run the same policy yet disagree on a token's logprob due to floating-point non-associativity. SkyRL's IsoExec combines an execution contract with a unified model, aligning rounding-sensitive execution choices across rollout and training. Bitwise parity holds across different TP, EP, and SP layouts. For Gated DeltaNet, the chunkwise-parallel recurrent algorithm makes parallel training and prefill bitwise identical to recurrent decode. Qwen3.5-35B-A3B, DAPO, 8xH100, 50 steps: logprob diff 1.6e-2 to 6.7e-7, full-step overhead 25.3% ✅ vLLM's scheduler and CUDA graphs still apply. Thanks to @JiangAlexander1 and the SkyRL team at @NovaSkyAI. 🔗
Show more
Want to RL-train a model on your own task without building the infra? Use our new cookbook with @hud_evals: → Define your task + grader in HUD → Sample + train the model on Fireworks Define the task once. Train and evaluate against the same environment.
Show more
How auditing RL Environments (ie. looking at the data) will help us explain model behavior "models are benchmark shaped" bc GRPO style RL is essentially - synthetic data generation by a model - induced by the environments + harness design - behavior (trajectories) reinforced by verifier design the "data" is RL environments & tasks, a model generates a series of tokens in its exploration of that environment trying to solve the task we don't really know what it will do, but we reward behavior that aligns with the verifier Encouraged behaviors: important: the only behavior that even has an opportunity to get rewarded is trajectories our environment "induces" or "encourages". something that isn't produced, can't be rewarded if our Harness (via a tool description) or Task design encourages lots of parallel search calls over a massive document corpus to find a few facts, the model will start copying that behavior because it got rewarded for it in the future, when some other problem is roughly shaped like that search Task, the model will mimic similar behaviors whether or not that strategy is optimal models do things that their data encouraged them to do via reward Tracing Behavior to Environments: every behavior can be fuzzily traced by the accumulation of different environment design and reward combinations the model saw over training i think we all explicitly know this but we don't have the interpretability tooling to understand induced behaviors at a fine-grained level over a massive training run this is why I'm very exicting that ppl are swarming around Trace analysis over Evals/Environments I bet a lot of "weird" or "bad" behavior comes from some quirks in environment design that inadvertantly rewarded that behavior the only thing that even has an opportunity to get rewarded Vestigial Behavior RL is in part inherently exploratory. Models need to do a bunch of stuff in rollouts to figure out what works, the verifier will reward what works But the verifier won't explicitly penalize any behavior that has a neutral effect but that co-occurs in a successful rollout this is how we get behavior that used to be helpful in some tasks, doesn't hurt current tasks, and thus persists over time similar to how humans have vestigial organs like the appendix i bet we can collectively figure out a ton about model behavior by mechanistically working backwards from the data and I'm bullish humans and agents working together will make a lot of progress on this in the next year
Show more