Register and share your invite link to earn from video plays and referrals.

Search results for Hy4Preview
Hy4Preview community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Hy4Preview
Hy4 Preview from @TencentHunyuan is live on AI Gateway. • 𝚝𝚎𝚗𝚌𝚎𝚗𝚝/𝚑𝚢𝟺-𝚙𝚛𝚎𝚟𝚒𝚎𝚠 • Open source MoE at 770B params, 49B active, 1M context window
Show more
Hy4 preview is now available in OpenCode Go 770B/49B · 1M context · built for coding agents
0
71
1.7K
35
Forward to community
Hy4 Preview def wins this 3D sim of the ISS in orbit contest! the best frontend output in comparison to Qwen 3.8 Flash, GLM 5.3, and DeepSeek V4 Flash Vision what a crazy week, and GLM 5.3 open weight is coming tmr 🤯 @TencentHunyuan
Show more
📢 Hy4 Preview Is Now Live on API! Developed by Tencent Hunyuan, Hy4 Preview is a 770B sparse MoE open-weights model activating 49B parameters per token, demonstrating strong performance (2.99/4.00) in internal expert evaluations on 203 engineering tasks. Official API access is now available! 👉 Try now: 🔗 Learn more:
Show more
Tried Hy4 preview through WorkBuddy and gave it a pretty vague prompt for a first person roller coaster ride, front row seat, make it feel fast, that was basically it. No detailed spec, no reference video, just a rough idea in a sentence. What came back actually held up. Steep drop that felt like the bottom dropping out, banked turns that leaned the right way, scenery rushing by close enough that you register the speed instead of just watching it happen. For something built off a loose prompt like that, I did not expect it to nail the feel this well. Coding and 3D rendering have clearly gotten a real upgrade here, and this took one pass to look this good. Free to try in @WorkBuddy_AI for the next two weeks if you want to mess with it yourself. @TencentHunyuan @TencentAI_News
Show more
Tencent Hy4 Preview leads on SWE-bench Pro. 770B, 49B active, 1M context, and their biggest generational leap measured to date. It’s exciting to see another open weights model compete against the frontier. Try in Cline with: 1. npm i -g cline 2. /model 3. Select Hy4 preview
Show more
🔥 @TencentHunyuan Hy4 preview just cited a Zhihu post with a surprisingly simple challenge to DeepSeek’s mHC: What if the doubly stochastic matrix is overcomplicating things—and Identity actually works better? Today, let’s revisit the cited post, “Your DeepSeek mHC May Not Need the ‘m’,” by Zhihu contributor 涮月亮的谪仙人. After training Qwen3 1.7B and 8B dense models from scratch on 150B tokens, the experiments found: 💡 Identity HC > mHC > mHC-lite > orthogonal mHC In other words: simply setting H_res = Identity beat the Sinkhorn-Knopp constrained version. Here’s why 👇 1️⃣ What does mHC actually change? Standard Transformers have one residual stream. Hyper-Connections (HC) expand it to multiple parallel streams: 🔹 H_pre reads from the streams 🔹 H_post writes back to them 🔹 H_res mixes information between them DeepSeek’s mHC constrains H_res to a doubly stochastic matrix via Sinkhorn-Knopp, helping preserve norms and stabilize propagation. But the experiments suggest a simpler question: Do we need H_res mixing at all? 2️⃣ mHC seems to learn something close to Identity anyway For a single layer, the learned H_res is already close to Identity: diagonal ≈ 0.96, off-diagonal ≈ 0.01 But multiply H_res across many layers, and it gradually collapses toward a uniform 0.25 matrix. So each layer may look almost like Identity, while their cumulative effect becomes uniform mixing. The simplest fix? Just set H_res = I. 3️⃣ Identity preserves stream semantics With Identity, each residual stream stays where it is: Stream 0 stays Stream 0. Stream 1 stays Stream 1. No repeated reshuffling, no cumulative mixing, and Iᴸ = I. H_pre and H_post also no longer need to track where each stream has been repeatedly moved—they simply learn where to read and where to write. 4️⃣ But cross-stream communication still happens Setting H_res = I does not isolate the streams. The projection that generates H_pre and H_post already sees all residual streams, making H_pre input-dependent. So information can still be dynamically aggregated across streams before Attention/MLP and written back afterward. H_res isn’t the only mechanism for cross-stream interaction. 5️⃣ Why might Sinkhorn hurt? Repeated products of positive doubly stochastic matrices tend toward uniform mixing. In the Qwen3-1.7B experiment, after 56 HC modules, the minimum singular value of the accumulated H_res product reached just: 9.2 × 10⁻¹⁸ By ~10 layers, the four streams were already approaching the same 0.25 uniform mixture. Sinkhorn also comes with extra cost: 20 iterations, backward recomputation, extra parameters, and approximation error. Identity has none of these—and preserves the residual signal exactly. 6️⃣ More sophisticated alternatives didn’t win either The experiments also tested mHC-lite, softmax-weighted convex combinations, and orthogonal variants using Cayley/Givens transforms. The observed ranking remained: Identity HC > mHC > mHC-lite > orthogonal mHC The simplest design won. 💡 The takeaway DeepSeek’s mHC uses sophisticated manifold constraints to stabilize Hyper-Connections. But these experiments suggest that H_res itself may not need to be learned or mixed at all. Sometimes the best manifold constraint is the most boring one: H_res = I. Or, as the original post puts it: Maybe DeepSeek’s mHC doesn’t need the “m.” 😆 👉 Read the full Zhihu post for the training curves, mathematical analysis, H_res visualizations, and implementation details: #DeepSeek# #mHC# #Hunyuan# #Tencent# #LLM# #Transformer# #AIResearch# #AI#
Show more
DAILY SITUATION RECAP: - Anthropic wins in federal court - South Korea is building a free AI service for its citizens - Tencent releases Hy4 Preview - Trump admin is trying to close China's AI chip loophole Federal court rules in favor of Anthropic. A federal judge ruled that the Pentagon acted illegally by labelling Anthropic a supply chain risk. U.S. District Judge Rita Lin said that Anthropic was being illegally punished for criticizing the Department of War’s views on how they would use Anthropic’s models internally, and who had the final say for what was allowed. Anthropic sued the Pentagon in March after they were designated a supply chain risk that same month. a16z announces $1.1 billion Machine Age Fund, which will fund the physical buildout of AI hardware, including chips, networking, memory, and storage. a16z notes that all layers of the AI stack are hitting a wall of supply chain capabilities and that innovation and investment is needed to meet demand. Trump admin working on rule to stop China’s remote access to chips. The PRC currently has remote access to AI computing clusters in nearby countries like Thailand and Singapore. This access is seen as a loophole in current export bans of advanced AI chips to China. The Department of Commerce is likely to share this rule with trade groups in September. Memory maker CXMT’s first-half revenue increases 873.64%. CXMT posted a net profit of 77.6 billion yuan ($11.5 billion USD), a dramatic turnaround from a net loss of 2.3 billion yuan ($342 million USD) in 2025. Tencent releases Hy4 Preview, an open-source 770B parameter model with a 1M token context window. Tencent says the model is a frontier open-source model built for productivity. South Korea is building a free, nationwide AI service citing a need to strengthen the nation’s domestic AI options for its citizens. The government-backed program aims to provide free AI agents and chatbots to the public and has chosen SK Telecom, KT Corporation, and Kakao to build these services. Anthropic’s talks to acquire MatX for $7 billion abandoned. MatX is a semiconductor company that makes custom chips designed to support AI models. The talks with MatX show Anthropic’s continued interest to design custom chips for their models. Anthropic plans to spend $36 billion on Google’s AI chips in the future. The X Safety team found bot farm manipulating debate over American AI policy. The investigation concluded that a bot farm of 200 accounts was making posts claiming that “AI data centers are driving up household electricity prices and straining the grid”
Show more
On behalf of @NousResearch since they haven't said it yet (yall slacking over there) Tencent Hy4 preview is available on Nous Portal for 20% off.
A new open weight Chinese model has hit the timeline!! This time it’s from Tencent’s Hunyuan team with Hy4 preview. It’s a massive MoE with 770B total parameters / 49B active per token, a 1M context window, and the weights are released under Apache 2.0. Important caveat - Tencent says this is still an early Hy4 checkpoint with more pretraining + post training to come. I actually like how honest their benchmark sheet is. They show plenty of places where the model is still behind despite its size. Hy4 gets 85.4 on Terminal Bench 2.1, basically right in the frontier cluster, and jumps from Hy3’s 28.0 -> 64.3 on DeepSWE. But it’s still behind Kimi K3 at 74.0 and Claude Opus 5 at 74.7 there. On ProgramBench it gets 17.5 vs Claude’s 39.5, SWE Atlas Refactoring 53.3 vs 60.0, and Humanity’s Last Exam 43.4 vs 53.2. It’s a 49B active open weight model that is competitive for its size on a bunch of hard coding/agent benchmark. One thing Tencent also reported on GitHub is they ran a 163 person internal blind eval across 203 engineering tasks where Hy4 slightly beat GLM 5.3 and Kimi K3, which is pretty interesting.
Show more