Register and share your invite link to earn from video plays and referrals.

Search results for 3)✅
3)✅ community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 3)✅
Season 13 - Season 3 ✅ With the upcoming Beachside skin in #Overwatch#, Juno has received at least one skin every season (beside her debut one) and is on a 17-skin streak! 🫣
Show more
0
37
4.2K
221
Forward to community
Three fresh models just dropped on LobeHub 👀 ✅GLM-5.3 ✅Gemini 3.7 Flash ✅Grok 4.6 Quick vibe check: GLM-5.3 → China’s latest open-source SOTA, with a big post-training upgrade over GLM-5.2. Gemini 3.7 Flash → a serious jump from 3.6 Flash. Already looking very good for frontend + game dev. Grok 4.6 → same base as 4.5, but noticeably sharper writing and responses. All live now. Pick one. Build something weird.
Show more
Anti-AFD rally in Berlin today: ❌ German flags (0) ✅ LGBTQ flags (3) ✅ Antifa flags (2) This says everything you need to know
0
315
36.6K
2.9K
Forward to community
this is wild. 3rd time in a row i’ve handed Claude a pile of transcripts, emails, and data of some situation and asked to guess if i'll get a yes or a no in some days based on that. 3/3 ✅
In RL training, a vLLM rollout engine and a Megatron trainer can run the same policy yet disagree on a token's logprob due to floating-point non-associativity. SkyRL's IsoExec combines an execution contract with a unified model, aligning rounding-sensitive execution choices across rollout and training. Bitwise parity holds across different TP, EP, and SP layouts. For Gated DeltaNet, the chunkwise-parallel recurrent algorithm makes parallel training and prefill bitwise identical to recurrent decode. Qwen3.5-35B-A3B, DAPO, 8xH100, 50 steps: logprob diff 1.6e-2 to 6.7e-7, full-step overhead 25.3% ✅ vLLM's scheduler and CUDA graphs still apply. Thanks to @JiangAlexander1 and the SkyRL team at @NovaSkyAI. 🔗
Show more
DeepSeek V4.1 Gives Prefill and Decode Different Compute Paths Prompt tokens mostly traverse 20 layers; generated tokens traverse all 40. This Causal Encoder-Decoder design nearly halves long-input prefill while preserving autoregressive generation. Zhihu contributor 潜龙勿用, Changxin Ke(柯昌鑫), a graduate researcher at ICT, CAS, explains how its architecture and post-training were designed together. 1️⃣ A causal encoder, not T5 The 40-layer backbone is split into a 20-layer causal encoder and a 20-layer decoder. Both remain causal. The encoder processes the prompt and supplies the decoder’s global KV. Most prompt tokens avoid the decoder stack, while generated tokens still run through all 40 layers. That means 8B active parameters per prefill token versus 16B per decode token. 2️⃣ Most layers share global memory CSA2 uses three modes: Full creates global KV and an index; Reindex shares the KV but selects new positions; Reuse shares both. Only four layers create independent global KV, four reindex it, and 30 reuse both KV and the latest index. Each layer retains its own query and local SWA state. The first Full layer builds up to 16,384 candidates. Later layers search this pool for their Top-512. From 4K to 1M context, decode FLOPs per token rise only about 25%. 3️⃣ Serving approximations enter training Exact reconstruction of decoder SWA states would replay 2,560 prompt tokens. V4.1 replays only the final 128 encoder outputs and trains the model to tolerate the approximation. With FP4 global KV, cache falls to 890 bytes per token, persistent cache to roughly one eighth of V4-Flash, and prefill compute close to half. Bounded replay and constrained retrieval are not last-minute serving tricks. The model experiences them during post-training. 4️⃣ Post-training is an evolving Agent system Each RL task combines a problem, environment, and verifier. New trajectories can reveal shortcuts, broken environments, or verifier errors and send the task back for repair. V4.1 trains across multiple harnesses, merges checkpoints between RL runs, and finishes with on-policy distillation from more than 40 teachers. Raising reasoning effort from 25 to 100 increases output length about 2.5×, while average Pass@1 across eight benchmarks rises from 67.1% to 76.3%. ✅ The real design choice V4.1 aligns model structure, cache policy, retrieval limits, training environments, and inference around long-running Agents. The model learns under the same constraints the deployed system will actually impose. 🔗 Full analysis: #DeepSeek# #DeepSeekV41# #LLMArchitecture# #AIAgents# #LongContext# #AIInfra#
Show more
📊 𝐒𝐓𝐀𝐓: Arsenal’s Premier League opening-day record under Mikel Arteta: ✅ 3-0 vs Fulham ❌ 0-2 vs Brentford ✅ 2-0 vs Crystal Palace ✅ 2-1 vs Nottingham Forest ✅ 2-0 vs Wolves ✅ 1-0 vs Manchester United
Show more
Goals: 1) >17mph pace ✅ 2) finish in 6 hrs ✅ 3) Elapsed time < 7hr ✅ 4) Don’t have a heat stroke ✅ LFG
🚨 🚩 Live: The penalty shootout begins... Morocco Ismael Saibari scores ⚽️ Netherlands 🇳🇱 - 2 ✅ ❌ ✅ ❌ ❌ Morocco 🇲🇦 - 3 ❌ ✅ ✅ ❌ ✅