가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Lou
@louszbd
Code/Work with GLM
가입 December 2025
1.6K 팔로잉 중    21.9K 팬
We were bringing up the inference stack for GLM-5.3-Flash. And sitting there watching agent work, we found that what it got back after a change mattered about as much as how good the model was. That's what we mean by dense feedback. Our infra agent runs on GLM-5.3. It read through kernels tuned by hand and wrote down what it found as optimization skeletons. Whatever held up went back in so the next kernel takes less work. Early days, but it's already running in production. Different labs put self improvement in different place. Here's a small piece of ours from our everyday work.
더 보기
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.
더 보기