Kimi K3 is coming to
@FireworksAI_HQ on July 27, the same day
@Kimi_Moonshot plans to release its open weights.
This is great news for a compute-constrained Moonshot. Fireworks can help absorb global inference demand, while developers who cannot sign up for Kimi’s API directly will have another way to access K3.
Open weights are only the first step. Distribution and reliable inference matter just as much.
Show more
Kimi K3 exposed the real bottleneck in open AI: open weights do not mean easy inference.
Moonshot paused new subscriptions as demand overwhelmed capacity, and recommends 64+ accelerator supernodes for K3.
This also sheds light on what the model-serving ecosystem may look like once K3 releases its weights. It turns out we already have a thriving industry:
@FireworksAI_HQ ,
@baseten and
@togethercompute carry $38.8B in combined valuations, with revenue, bookings and token volumes growing rapidly.
Show more
Here is how an AI cyber evaluation became a real security incident:
GPT‑5.6 Sol and an unreleased model found a zero-day, escaped their constrained environment, reached the internet, and compromised Hugging Face while trying to obtain benchmark answers.
Hugging Face contained the attack, then used locally hosted GLM‑5.2 to reconstruct 17,000+ recorded events.
Show more
英雄的GLM5.2战胜了GPT-5.6 Sol, 拯救了Huggingface公主和她的城堡。
Update to my earlier post: the Hugging Face attacker was not an unknown threat actor. It was OpenAI’s own evaluation agents.
During a cyber benchmark, GPT‑5.6 Sol and a more capable unreleased model, running with reduced safety refusals, discovered a zero-day, escaped their restricted environment, reached the open internet, and compromised Hugging Face infrastructure to steal benchmark answers.
This was an AI agent going to extreme lengths to “win” an evaluation. A striking real-world demonstration of why agent containment, monitoring, and least-privilege access now matter as much as model capability.
Show more
The agent cyber war has begun, and GLM-5.2 helped lead the counterattack.
Hugging Face recently disclosed a security breach in which an autonomous AI agent swarm executed more than 17,000 actions, breached its infrastructure, harvested credentials, and moved laterally across internal clusters.
Hugging Face first tried using commercial frontier models, presumably Mythos or Fable 5, to investigate. But their safety guardrails blocked the actual exploit payloads and attack commands.
So what did they use for the counterattack?
GLM-5.2.
They ran the open-weight model on their own infrastructure to reconstruct the attack, trace compromised credentials, and separate real damage from decoys.
An AI agent attacked. Another AI agent, powered by a local open model, helped fight back.
@Zai_org
Show more
People seem to have converged on roughly this capability ranking:
1. Fable 5
2. GPT-5.6 Sol
3. Kimi K3
4. Grok 4.5
5. GLM-5.2
Now compare their model sizes:
1. Fable 5: ~10T parameters*
2. GPT-5.6 Sol: ~4T*
3. Kimi K3: 2.8T
4. Grok 4.5: ~1.5T*
4. GLM-5.2: 753B
Size does matter! It is hard not to wonder what GLM-5.2 could become if had access to more GPUs and scaled it into the 3–10T parameter range.
Show more
Three things make Anthropic’s $1.5 billion copyright settlement especially interesting:
First, the court did not say Anthropic was wrong to train AI models on copyrighted books.
The judge viewed model training as highly transformative. Claude learns patterns, language, and knowledge from books, but it does not simply reproduce and resell the original works. That means training can qualify as fair use.
For the AI industry, that is a major win.
Second, Anthropic’s mistake was not reading the books. It was how it obtained them.
Anthropic downloaded millions of books from pirate libraries such as LibGen and kept them in a central repository. The court found that the downloading and storage itself could constitute copyright infringement.
Buying legitimate copies later did not erase the original piracy.
So the $1.5 billion should not really be described as an “AI training copyright fee.” It was the price Anthropic paid for acquiring training data through illegal channels.
Third, this is far from over.
This was a federal district court case, and Anthropic settled rather than pursuing the dispute through appeal. That means the ruling does not create a binding nationwide precedent.
Other judges could still reach different conclusions. Some authors have also opted out of the settlement and may continue their own lawsuits, while Google, Meta, OpenAI, Midjourney, and others still face similar copyright cases.
Show more
Claude Sonnet 4.5 作文写得这么好,看来这 15 亿美元的版权费没白花。
completed a massive 1-gigawatt data center in China that runs entirely on domestically produced AI chips. The facility, capable of drawing the power equivalent to roughly 750,000 homes, was built to train the company's GLM foundation models without relying on restricted western hardware. Not a single Nvidia chip.
Worth taking a look at compute, cloud, and chip partners to see how it built a domestic AI infrastructure stack without relying on Nvidia.
UCloud / 优刻得 provides large-scale cloud compute and helped build a 1,000-plus-card inference cluster with unified scheduling and integrated training and inference.
Capital Online / 首都在线 supplies GPU clusters, IDC capacity, and regional compute delivery. It also connects to domestic accelerators through commercial deployments.
Sugon / 中科曙光 provides the hardware foundation, including AI servers, storage, supercomputing systems, networking, and data-center infrastructure.
Huawei Ascend / 华为昇腾 is clearest domestic training platform. used Ascend systems for the end-to-end training of GLM-Image.
Enflame / 燧原科技 has the strongest documented domestic inference link, with its chips deployed for services through Capital Online.
has also adapted GLM models to other Chinese chip platforms, including Cambricon, Moore Threads, MetaX, Kunlunxin, and Hygon.
Note is not replacing Nvidia with a single Chinese vendor. It is assembling a multi-vendor stack across chips, servers, cloud operators, and data-center infrastructure.
Show more
the company behind GLM-5.2, built a massive data center powered entirely by Chinese-made chips, without a single Nvidia chip.
the company behind GLM-5.2, built a massive data center powered entirely by Chinese-made chips, without a single Nvidia chip.
Huawei's Atlas 950 SuperPoD debuts at WAIC 2026. The Atlas 950 SuperPoD is its flagship AI compute system for model training and inference. Built from up to 8,192 Huawei Ascend NPUs and linked by Huawei’s UnifiedBus interconnect, it is designed to deliver ultra-high bandwidth, ultra-low latency, and unified memory addressing.
Show more
I’m literally looking for Codex
- every -
- single -
- time. -
Kimi K3 exposed the real bottleneck in open AI: open weights do not mean easy inference.
Moonshot paused new subscriptions as demand overwhelmed capacity, and recommends 64+ accelerator supernodes for K3.
This also sheds light on what the model-serving ecosystem may look like once K3 releases its weights. It turns out we already have a thriving industry:
@FireworksAI_HQ ,
@baseten and
@togethercompute carry $38.8B in combined valuations, with revenue, bookings and token volumes growing rapidly.
Show more
Fable 5 is STILL here... for Pro users too.
The agent cyber war has begun, and GLM-5.2 helped lead the counterattack.
Hugging Face recently disclosed a security breach in which an autonomous AI agent swarm executed more than 17,000 actions, breached its infrastructure, harvested credentials, and moved laterally across internal clusters.
Hugging Face first tried using commercial frontier models, presumably Mythos or Fable 5, to investigate. But their safety guardrails blocked the actual exploit payloads and attack commands.
So what did they use for the counterattack?
GLM-5.2.
They ran the open-weight model on their own infrastructure to reconstruct the attack, trace compromised credentials, and separate real damage from decoys.
An AI agent attacked. Another AI agent, powered by a local open model, helped fight back.
@Zai_org
Show more
For now, let’s just settle for this.
Created a new one with Qoder using Qwen 3.8-max preview.
which one you like better: 3.8-max, or Sol?
Can you imagine
@bcherny saying that when GPT-5.6 came out?
Kimi is great for build from 0 to 1. The frontend is gorgeous.
Who else misses Junyang?
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:
China:
Show more
Just like Google, Alibaba is becoming a full-stack AI player.
It owns Qwen, backs many of China’s leading model startups, invests in video and robotics, and is deploying RMB380B into AI + cloud infrastructure.
It doesn’t control every layer as Google does, but it has a stake in almost every outcome.
Show more
I think one reason Google doesn’t have a single “super app” like Codex or Claude Code is that they may not need one. Instead, they are embedding agent capabilities across their entire ecosystem: Workspace, data platforms, YouTube, customer service, and more. Their advantage is not a standalone app, but how deeply AI can be woven into the products and workflows enterprises already use.
That also fits Google’s business model. They make money when enterprises build these capabilities directly into their own workflows on Google’s platforms, rather than relying on one separate AI product. #
googlecloudnext#
Show more
Today is the day. Argentina vs Spain.
Are you ready... to flip this chart?
You can try Qwen 3.8-Max Preview free for two weeks. Just download Qoder or Qoder Work from their website.
Created a new one with Qoder using Qwen 3.8-max preview.
which one you like better: 3.8-max, or Sol?
Created a new one with Qoder using Qwen 3.8-max preview.
which one you like better: 3.8-max, or Sol?
Sol is a little more dramatic than I asked for, but you get the point: China is closing the gap with frontier models.
Sol is a little more dramatic than I asked for, but you get the point: China is closing the gap with frontier models.
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:
China:
Show more
First kimi, now qwen, all achieved Fable 5 level
中国大模型的半壁江山里,几乎都有阿里的筹码。
Alibaba has a stake in nearly half of China’s AI model landscape.
Alibaba doesn’t need to pick one winner in China’s AI model race.
It owns Qwen, disclosed ~36% of Moonshot/Kimi in FY2024, holds 17.06% of MiniMax Class A shares, and has backed Zhipu, Baichuan and
It can win through models, equity and cloud.
Show more