Register and share your invite link to earn from video plays and referrals.

Search results for AIChips
AIChips community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including AIChips
⚙ How GLM-5.3-Flash Served 70 Trillion Free Tokens on Chinese Chips GLM-5.3-Flash's anonymous "Ox Alpha" trial burned through roughly 70 trillion tokens in a week — and Zhipu says all of it ran on Chinese AI chip clusters. For many observers, that is a bigger story than the model itself. Zhihu contributor 恋猫 breaks down the systems engineering that made it work. His framing: the model architecture first reduces how much data needs to move, then the inference system compresses what remains. The reported scale is around 100,000 Chinese chips — from Huawei, Moore Threads and Hygon, according to @latepostnews — though @Zai_org itself has only said "tens of thousands." 1️⃣ The architecture cuts data movement first Versus GLM-5.3, Flash lowers attention compute by 3.01x and KV cache by 4.44x. Less data shuttling between memory and compute is the foundation everything else builds on. 2️⃣ Compute-for-bandwidth: trading FLOPs for HBM relief The bottleneck on these chips is HBM: the compute units still have headroom while memory bandwidth is nearly saturated. The fix, loosely speaking, turns "every step: read the full state, write the full state" into "read the state, recompute a little, write back periodically." A bit of extra matrix math buys a large reduction in HBM write traffic. 3️⃣ Communication-for-bandwidth: shard the cache across cards The cluster also uses high-speed inter-chip links and aggregated bandwidth to cut how much data each card must keep resident. The author's example: rank 0 holds the KV and indexer cache for some layers, rank 1 holds the rest. No card stores every layer's cache long-term; data is prefetched over the interconnect as each layer executes. The trade-off is real — more communication, far less resident cache per card. 4️⃣ EPD separation: three pools, scaled independently Finally, Encode, Prefill and Decode are split into three independently scalable resource pools, which matters most for multimodal traffic. Together, these pieces are what let the cluster run at high utilization. 5️⃣ The result: 3x end-to-end, cost "comparable to NVIDIA" Zhipu's own claim: end-to-end serving performance improved 3x on the same hardware, bringing per-token cost close to mainstream NVIDIA GPUs. The author's reading: individual Chinese cards may still be weaker, but model-system co-design lets the cluster as a whole reach international-mainstream throughput and cost. He adds a wry footnote: Zhipu's API used to be notorious among Chinese developers for 429 rate-limit errors. That it could absorb this launch's traffic at all — entirely on Chinese silicon — is, in his words, proof that optimization never ends. 🔗 Full Reading: #GLM# #Zhipu# #AIChips# #AIInfra# #LLM# #Inference# #OpenWeights#
Show more
💥 GLM-5.3-Flash: Reported Opus 4.8-Level Scores at One-Tenth the Price Zhipu has confirmed that Ox Alpha — the anonymous model that dominated OpenRouter and OpenCode last week — is GLM-5.3-Flash, a 320B-parameter MoE with 18B active, released with open weights and an API at roughly one-tenth of GLM-5.3's price. Zhihu contributor 小小将 opens with a confession: he had hoped Ox Alpha was an external team fine-tuning the open GLM weights. He was wrong — it was Zhipu's own model all along. His broader take: at this performance and this price, Flash becomes the new "kill line" for large models — the bar below which rivals simply get priced out. 1️⃣ The numbers behind the "kill line" Flash is somewhat larger than DeepSeek-V4 Flash, at 320B total and 18B active. On Zhipu's own benchmark comparisons, overall capability is roughly level with Claude Opus 4.8. More reference points he cites: 🔹 57 on the Artificial Analysis Intelligence Index — around GPT-5.6 Terra's level, slightly below GLM-5.3. 🔹 No.5 on Code Arena, one spot above GLM-5.3. 🔹 API at 1/10 of GLM-5.3's price, with a limited-time half-off promo bringing it to 1/20. 2️⃣ Same DeepSWE score, a fraction of the cost The author's sharpest comparison is on DeepSWE, where Flash scores 63% at a single-task cost of $0.24. DeepSeek-V4-Pro hits the same score at $1.67 per task — roughly seven times the cost for equivalent results. This is the author's cost arithmetic on reported figures, not an independent measurement. 3️⃣ The architecture that cuts the bill Part of the price drop is structural. Flash uses a hybrid of linear and sparse attention, sharply cutting attention compute, and a new IndexPool that compresses the indexer's cache from four copies to one — reducing latency and memory at million-token context. Versus GLM-5.3, the company reports attention compute down to about 1/3 and KV cache down to about 1/4.4. 4️⃣ The compute mystery, answered by domestic chips One reason few believed Ox Alpha was Zhipu's: the company was not thought to have enough spare compute for a massive free public test. The answer, per Zhipu: Flash is served from domestic Chinese chip clusters. The team built a custom inference engine on SGLang and worked around limited VRAM and bandwidth with quantization, layered deployment, and trading compute for communication — lifting end-to-end serving performance about 3x on the same hardware, with per-token cost now close to mainstream NVIDIA GPUs. 5️⃣ Chinese silicon just passed its biggest stress test The author's conclusion: Chinese chip clusters have shown they can carry large-scale inference for a frontier-level model — including a free, record-breaking public trial. If that holds, he argues, the outlook for Chinese models just got a lot brighter. 🔗 Full Reading: #GLM# #Zhipu# #GLM5Flash# #AIChips# #LLM# #AIInfra# #OpenWeights#
Show more
🔌 GLM-5.3-Flash Served Its Viral Debut Entirely on Domestic Chinese Chips Zhipu's GLM-5.3-Flash — the 320B-A18B model revealed this week as the anonymous "Ox Alpha" — set usage records on OpenRouter and OpenCode during its undercover test. The company says all of that traffic was served by domestic Chinese chip clusters. Zhihu contributor 刘延 reconstructs how Zhipu lined up this infrastructure, and reads the official engineering details for hints about which chips are actually doing the work. The core judgment: the bigger story here is not the model itself, but that a frontier-level model handled global-scale, real-world inference on domestic silicon. 1️⃣ The timeline behind the launch The author pieces together a sequence from public reporting. 🔹 Zhipu was reported to have acquired an infrastructure company. 🔹 LatePost reported Zhipu had brought 50,000 domestic cards online; around the same time, its CodePlan subscription got cheaper with generous bonus quotas. 🔹 Ox Alpha went live anonymously and, in the author's words, blew up worldwide. 🔹 Zhipu then confirmed every request in that test ran on domestic chips. 🔹 The latest LatePost report puts the deployment at 100,000 domestic cards. Note the card counts come from media reports, not Zhipu itself. 2️⃣ The engineering: surviving 1M context on constrained hardware Zhipu's own statement is unusually specific about the constraints. The main bottleneck on these chips is memory capacity and bandwidth, and supporting a 1M-token context is the hardest part. The company's listed optimizations include trading compute for bandwidth and communication for memory, intra-node tensor parallelism for the linear attention and the LM head, ReplaySSM, W8A8 quantization, INT8/FP8/BF16 mixed cache quantization, and Layer Split. 3️⃣ Which chips? Reading the precision hints Here the author speculates, and it should be read as inference, not confirmation. 🔹 FP8 support suggests Moore Threads could be handling prefill, or possibly Hygon's DCU-3. 🔹 INT8 points toward Ascend 910B/C as the likely backbone. 🔹 No mention of FP4 suggests the newer Ascend 950 is probably not in the mix. 4️⃣ Why this matters If the reporting holds, this is the first time domestic Chinese chip clusters have carried a frontier model's global production traffic at this scale — including a free, record-breaking stress test from developers worldwide. The author treats it as a proof point: China's domestic chips are no longer just for training experiments or internal pilots, but can serve a top-tier model to the open internet. 🔗 Key links: Official announcement: Open weights (MIT): 🔗 Full Reading: #GLM# #Zhipu# #AIChips# #AIInfra# #Ascend# #Semiconductors# #OpenWeights#
Show more
AI’s biggest bottleneck is moving data and that could still create huge opportunities for optical networking companies (Save this) The chart shows a 1.6T optical transceiver, a device that transfers data between AI servers, switches, GPUs, and fiber optic networks. 1.6T means it can theoretically move up to 1.6 terabits of data per second, or 1,600 gigabits and that is about twice the speed of an 800G connection. This technology is important because AI data centers contain thousands of GPUs that must constantly exchange information. As AI models become larger, slow connections can leave expensive processors waiting for data but faster optical links help reduce that bottleneck and allow AI clusters to operate more efficiently. This image shows two directions of travel. The TX path converts electrical data from a server or switch into light which travels through fiber. The RX path receives that light and converts it back into an electrical signal for another device. And inside the module are several key components. Optical DSPs process and correct the signal, laser drivers control the lasers, modulators place data onto the light, and photodiodes convert incoming light back into electricity. Amplifiers, timing chips, thermal sensors, power management devices, and high speed connectors help the system operate reliably. Optical fiber becomes more attractive at higher speeds because copper connections lose efficiency over longer distances. At 1.6T, copper may only be practical across very short distances while optical technology can move data farther with better bandwidth and lower signal loss. This creates an investment opportunity beyond the companies making AI chips. NVIDIA remains a major beneficiary because its AI systems require fast connections between GPUs. Broadcom and Marvell could benefit from their networking chips, custom silicon, and optical connectivity products. Coherent and Lumentum are important optical suppliers with exposure to lasers, photonics, and high speed transceivers while Applied Optoelectronics is a more direct transceiver play and has announced a volume order for 1.6T data center products. Arista Networks and Cisco could benefit by selling the switches and networking systems that connect AI servers. Chinese suppliers such as Innolight, Eoptolink and Accelink Technology could also benefit as China expands its AI data center infrastructure. If you enjoyed reading this, make sure to follow @MelvinInvests for more photonics, AI infrastructure and semiconductor insights and turn on post notifications so you don't miss a single update. If you want to see exactly what I'm buying as an analyst at Milk Road Pro, check out the link below:
Show more
The AI bull run is just getting started and here is how you want to position before the biggest spending wave (Save this). The chart shows hyperscaler capital spending rising from $491 billion in 2025 to an estimated $950 billion in 2026 and $1.4 trillion in 2027 and by 2030, spending could reach approximately $3 trillion. The right side of the chart is especially important because analysts have continued raising their estimates. The 2026 forecast increased from $731 billion to $950 billion, while the 2027 estimate rose from $833 billion to $1.4 trillion which suggests the AI infrastructure buildout is happening faster and at a larger scale than previously expected. Now here is how you can benefit from all of this. Nvidia is the obvious beneficiary but investors should also look at the companies supplying the less visible parts of the AI ecosystem. Credo Technology makes high speed connectivity products that allow AI chips, servers, and switches to communicate while Astera Labs provides connectivity solutions that link CPUs, GPUs, memory, and storage inside AI servers. As AI clusters become larger, these companies could benefit from the need to move data faster between processors. Celestica manufactures and integrates servers, networking systems, and other data center hardware for large technology customers while Fabrinet produces complex optical and electronic equipment for other companies. These businesses may benefit as hyperscalers outsource more of the manufacturing required to build AI infrastructure. Applied Optoelectronics is a more direct optical networking play and it has announced a volume order for 1.6T data-center transceivers, which are designed to move data between next-generation AI systems. Coherent and Lumentum also provide lasers, photonics, and optical components used in high-speed networks. Semtech supplies signal conditioning and connectivity technology that helps preserve data quality as transmission speeds increase while Arista Networks and Cisco could benefit from selling the switches and networking systems that connect AI servers across data centers. Marvell is exposed to custom AI chips, networking, and optical connectivity and ass cloud companies develop their own AI processors, Marvell could benefit from helping them design and connect those systems. The power side of the buildout could create another group of winners. Advanced Energy Industries supplies power conversion systems used in data centers and semiconductor equipment while Modine provides thermal management products, while Vertiv supplies cooling, power, and data center infrastructure. This is exactly why we’re positioned across the entire AI infrastructure stack at Milk Road, not just there big names. If you want to see the trades we’re making around this spending wave, join us using this link.
Show more
$NVDA AI SERVER PRICES SET TO RISE 15%+ Some Nvidia customers have been notified that AI server prices will rise more than 15% in many cases for systems shipping early next year, including Vera Rubin and Grace Blackwell. Bloomberg says the increases will vary by chip generation and memory configuration, with soaring DRAM costs from Samsung, SK Hynix and Micron driving much of the pressure. Server makers supplying hyperscalers including Microsoft, Google, and Microsoft have already begun notifying customers. Amazon, Microsoft, Google and Meta are developing in-house AI chips, but remain heavily dependent on Nvidia for their data center buildouts.
Show more
⚡ AI Brief ⚡ Jane Street first tested Etched’s AI chips, then became its first paying customer. Now, the quantitative trading firm is leading Etched’s latest $700 million funding round at a $21 billion valuation. The round also included Kleiner Perkins, Sequoia, a16z, Tiger Global, Bain Capital Ventures, Blackstone, Stripes, Neo, Primary, and Positive Sum. Etched has now raised approximately $1.9 billion in total. Etched is building AI chips and full-scale inference clusters designed specifically for running AI models after they have been trained. Its technology focuses on increasing compute density and pooling memory across an entire rack, with the goal of delivering faster inference at a lower cost per watt. The company’s valuation has risen rapidly—from $5 billion in December, to $10.3 billion in July, and now $21 billion. Jane Street’s move from customer to lead investor also highlights growing investor interest in the infrastructure powering AI inference, not just the models themselves. To learn more about Etched, check out our Deep Dive Report via the link in bio. Sources: Etched, The Wall Street Journal Image Credit: Etched
Show more
Weekly AI Update The AI chip gold rush has a new banker worth $2.4B Volta Infra, a brand new AI cloud startup, just raised $300 million plus $5 billion in financing to help more companies afford Nvidia's pricey chips, not just the tech giants. Backed by Nvidia, Michael Dell, Andreessen Horowitz, and Altimeter, it has already locked in a $10 billion contract to supply cloud capacity to an unnamed AI developer over six years. The bet is that whoever solves AI's financing problem wins, though critics warn of risky circular deals and a coming shakeout. Google's $1.5B play for Mechanize could finally fix its coding AI problem Google is reportedly in talks to pay around $1.5 billion to license technology from Mechanize, a young AI startup building simulated environments and benchmarks to train coding agents, while also hiring some of its top evaluation experts. The move fits a familiar Google playbook of striking talent and licensing deals that sidestep antitrust scrutiny, much like its earlier tie-ups with Windsurf and The goal is clear: close the gap with Anthropic's Claude Code and OpenAI's Codex, which have pulled ahead in the fast growing market for AI coding tools. Anthropic wants to build its own chips to run Claude faster and cheaper Anthropic confirmed it is assembling a custom silicon team to design its own AI chips, hiring engineers who will develop hardware and models together so Claude can run faster and at greater scale. The move puts it alongside OpenAI, Google, and Meta, all of which have pushed into custom chips, though building an advanced one can cost close to half a billion dollars. It is the latest piece of a massive infrastructure push that also includes a $15 billion Texas data center campus and deals with $AMZN, $NVDA, $AMD, and Samsung.
Show more
Weekly AI Update ANTHROPIC EXPORT CONTROLS LIFTED The U.S. Commerce Department lifted export controls on Anthropic's Claude Fable 5 and Mythos 5, restoring global access to Fable 5 and expanding Mythos 5 availability to approved partners. The move ends a regulatory standoff and removes a potential competitive advantage for Chinese AI developers. META EXPANDS INTO AI CLOUD $META is developing a cloud infrastructure business to sell AI computing power and hosted AI models, positioning itself against AWS, Google Cloud, Microsoft Azure, and $CRWV. The initiative could monetize excess AI infrastructure, helping offset the company's massive investments in data centers and AI chips. MGX CLOSES $49B AI FUND Abu Dhabi-based MGX has closed a $49 billion AI investment fund, one of the largest ever, exceeding its $45 billion target. The fund, a major backer of OpenAI, Anthropic, and xAI, will invest across semiconductors, AI infrastructure, and AI platforms, underscoring continued global capital inflows into the AI sector.
Show more
Microsoft’s self-developed AI chips will be in mass production in 2027 as all work on design and capacity allocation remains smooth, media report, citing the head of Microsoft Taiwan, Sean Pien. Prior leaks had put production in the 2nd half of 2026. Pien said in the near term, the chips will power Microsoft’s internal needs, and won’t be open for third party use. He said the chips will drive down overall computing costs. $MSFT #semiconductors#
Show more