Register and share your invite link to earn from video plays and referrals.

Search results for 延松舞佳
延松舞佳 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 延松舞佳
🔌 GLM-5.3-Flash Served Its Viral Debut Entirely on Domestic Chinese Chips Zhipu's GLM-5.3-Flash — the 320B-A18B model revealed this week as the anonymous "Ox Alpha" — set usage records on OpenRouter and OpenCode during its undercover test. The company says all of that traffic was served by domestic Chinese chip clusters. Zhihu contributor 刘延 reconstructs how Zhipu lined up this infrastructure, and reads the official engineering details for hints about which chips are actually doing the work. The core judgment: the bigger story here is not the model itself, but that a frontier-level model handled global-scale, real-world inference on domestic silicon. 1️⃣ The timeline behind the launch The author pieces together a sequence from public reporting. 🔹 Zhipu was reported to have acquired an infrastructure company. 🔹 LatePost reported Zhipu had brought 50,000 domestic cards online; around the same time, its CodePlan subscription got cheaper with generous bonus quotas. 🔹 Ox Alpha went live anonymously and, in the author's words, blew up worldwide. 🔹 Zhipu then confirmed every request in that test ran on domestic chips. 🔹 The latest LatePost report puts the deployment at 100,000 domestic cards. Note the card counts come from media reports, not Zhipu itself. 2️⃣ The engineering: surviving 1M context on constrained hardware Zhipu's own statement is unusually specific about the constraints. The main bottleneck on these chips is memory capacity and bandwidth, and supporting a 1M-token context is the hardest part. The company's listed optimizations include trading compute for bandwidth and communication for memory, intra-node tensor parallelism for the linear attention and the LM head, ReplaySSM, W8A8 quantization, INT8/FP8/BF16 mixed cache quantization, and Layer Split. 3️⃣ Which chips? Reading the precision hints Here the author speculates, and it should be read as inference, not confirmation. 🔹 FP8 support suggests Moore Threads could be handling prefill, or possibly Hygon's DCU-3. 🔹 INT8 points toward Ascend 910B/C as the likely backbone. 🔹 No mention of FP4 suggests the newer Ascend 950 is probably not in the mix. 4️⃣ Why this matters If the reporting holds, this is the first time domestic Chinese chip clusters have carried a frontier model's global production traffic at this scale — including a free, record-breaking stress test from developers worldwide. The author treats it as a proof point: China's domestic chips are no longer just for training experiments or internal pilots, but can serve a top-tier model to the open internet. 🔗 Key links: Official announcement: Open weights (MIT): 🔗 Full Reading: #GLM# #Zhipu# #AIChips# #AIInfra# #Ascend# #Semiconductors# #OpenWeights#
Show more
💬 Customer Spotlight | Sytus Feed & Pinpoint "I was struggling with latency until Sytus Feed and Sytus Pinpoint transformed my trading. The ultra-low latency, rock-solid stability, and superior fill rates gave our team the ultimate execution edge." From solo traders to professional prop teams, achieving that ultimate execution edge is the goal. It’s about stability you can count on and fill rates that matter when every microsecond counts. Thank you to James for the trust and for the high recommendation to other trading teams. We remain dedicated to building the infrastructure that transforms trading performance. Interested in experiencing the difference? Feel free to reach out to us 📩 「在用 Sytus Feed 之前,我一直深受延遲之苦,它與 Sytus Pinpoint 徹底改變了我的交易。極致的低延遲、穩定性以及更高效的成交率,給了我們團隊更好的執行優勢。」 從獨立交易者到專業的自營交易團隊,追求終極的執行優勢是共同的目標。低延遲行情的價值在於可信賴的穩定性,以及在每一微秒都至關重要時,轉化為實質的成交績效。 感謝 James 對我們的信任,以及對其他交易團隊的大力推薦。我們將持續致力於打造更完美的基礎設施,協助交易者提升整體表現。 如果你也想體驗超低延遲行情所帶來的差異,歡迎與我們聯繫📩 #SytusFeed# #SytusPinpoint# #ExecutionInfrastructure# #LowLatency# #HighFrequencyTrading# #QuantTrading# #MarketMaking# #DigitalAssets# #QSG#
Show more