Register and share your invite link to earn from video plays and referrals.

Search results for megatron
megatron community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including megatron
Megatron and Blackarachnia have entered the review room… Things are about to get interesting. 🦖🕷️ #YOLOPARK# #beastwars# #megatron# #blackarachnia#
YESSSS🦖🕷️ AMK Series BEAST WARS Megatron & Blackarachnia Predacons Rise! 【Thanks for the photos by 筋肉方程式】 #YOLOPARK# #beastwars# #megatron# #blackarachnia#
Show more
In RL training, a vLLM rollout engine and a Megatron trainer can run the same policy yet disagree on a token's logprob due to floating-point non-associativity. SkyRL's IsoExec combines an execution contract with a unified model, aligning rounding-sensitive execution choices across rollout and training. Bitwise parity holds across different TP, EP, and SP layouts. For Gated DeltaNet, the chunkwise-parallel recurrent algorithm makes parallel training and prefill bitwise identical to recurrent decode. Qwen3.5-35B-A3B, DAPO, 8xH100, 50 steps: logprob diff 1.6e-2 to 6.7e-7, full-step overhead 25.3% ✅ vLLM's scheduler and CUDA graphs still apply. Thanks to @JiangAlexander1 and the SkyRL team at @NovaSkyAI. 🔗
Show more
On 10/10 she's live at the @theXtakeover, Tesla Giga Texas Nicki Minaj gets three random words on live TV and turns them into bars on the spot 🎤 "Six sides, that's a hexagon / I'm the big homie, Megatron / These girls can't see me like the Yeti" — @NICKIMINAJ 🎥 @FallonTonight (The Tonight Show Starring Jimmy Fallon, NBC · June 27, 2019)
Show more
Training and rollout logprobs matched bit for bit on ROCm. The @RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime. A 200-step Qwen3-8B GRPO run on 8× @AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides. Deep dive:
Show more