登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

思维怪怪
@0xLogicrw
写关于 AI 的一切 Shit Post @BeatingOfficial AI 信息流:
参加 May 2018
2.7K フォロー中    7.7K ファン
智谱披露,GLM-5.3 驱动的 Infra Agent 已经参与优化 GLM-5.3-Flash 的推理系统。从首次跑通到承接全部生产流量不到两周,端到端吞吐提升到最初的 3.2 倍。 联合创始人兼首席科学家唐杰在复盘中提到,Agent 现在最容易卡住的地方,往往不是不会写代码,而是不知道问题出在哪。比如一次修改让吞吐下降 20%,它知道自己改坏了,却不知道是哪一层出了问题,下一步该查什么。 于是智谱把资深工程师平时的排查方法做成一套 Agent 可以直接使用的工具。Agent 改完代码后,可以自己检查结果对不对、时间耗在哪里、哪种方案更快,再根据结果继续修改。 靠这套方法,Agent 找出了长上下文计算误差、KV 传输阻塞和解码 kernel 重复计算等问题。其中一个 kernel 重写后提速 1.71 倍。 唐杰判断,工程师的角色会越来越像「设计反馈的人」。人负责定目标、设边界和审核高风险修改,Agent 自己提假设、改代码、跑实验。离完整的 RSI 还很远,但「模型优化系统,系统再服务模型」这个最小闭环已经出现。
もっと見る
Two weeks. That's how long it took to go from GLM-5.3-Flash's first run on domestic accelerators to serving all of its production traffic, with 3.2× end-to-end throughput along the way. What I keep thinking about is who did much of the work: an Infra Agent powered by GLM-5.3. A model helping optimize the system that serves it. The conditions were hard. Limited memory and interconnect bandwidth. 1M-token context. Multimodal requests. An immature software stack where kernels were missing and documentation was often guesswork. Every optimization was a trade: compute for memory (ReplaySSM), communication for memory (intra-node tensor parallelism), precision for capacity (mixed INT8/FP8/BF16 caching), and disaggregation for scheduling freedom (Encode–Prefill–Decode). But the most important lesson wasn't about any single optimization. When the agent got stuck, it was rarely because it couldn't write the code. It was because it didn't know *why* things got worse. "Throughput down 20%" tells you something broke. It doesn't tell you which layer, which hypothesis, or what to test next. In RL terms, it's a sparse reward with a credit assignment problem. And an end-to-end benchmark that takes hours makes exploration painfully slow. Senior engineers solve this with an implicit process reward in their heads. They know when to check the timeline, when to run a microbenchmark, and which layer's output to compare. So we made that explicit. We call it dense feedback: layered verification interfaces the agent can call directly. Correctness feedback: did it compute right? System behavior feedback: where did the time go? Performance feedback: which option wins, under which conditions? Each signal has to be local, cheap, and objectively verifiable. Three things the agent found: First, precision drift in KDA's context-parallel path that grew with sequence length. The cause was TF32 rounding error compounding through chained state-matrix merges. The fix is now merged upstream in Flash Linear Attention (PR #1180#). Second, KV transfer never overlapped with DeepEP dispatch. The agent followed the call chain across the Python/C++ boundary and found that the intranode path never released the GIL. After the fix, transfer overhead fell from over 30% to under 1%. Third, a decode kernel recomputing the same normalization four times because of how it was chunked. The agent restructured it and got a 1.71× speedup. The idea came from "optimization skeletons" it had distilled by reading existing kernels across SGLang, FLA, and DeepGEMM. To be clear about the boundaries: humans still defined the goals, built the feedback environment, and reviewed every high-risk change. But the engineer's role is changing, from the person who solves the problem to the person who designs the feedback. There's a deeper implication too. A layered, verifiable feedback environment built on real infrastructure tasks is exactly what training the next generation of models needs most. Every task the agent completes can become training ground for its successor. We are still far from recursive self-improvement. But the smallest loop now exists. The model optimizes the system. The system serves the model.
もっと見る