注册并分享邀请链接,可获得视频播放与邀请奖励。

Wësche
@WescheNex1q
Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston
加入 January 2013
637 正在关注    2.4K 粉丝
GLM-5.3-Flash, 320B params, one DGX Spark, 256K context, 33.8 tok/s peak. Released 2 days ago Serving recipe with MTP speculative decode already up, incl. the branch pick that makes it 2x faster:
显示更多
GLM-5.3-Flash can now be run locally! ✨ Run 3-bit on 128GB RAM via Unsloth GGUF. GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks. Guide: GGUF:
显示更多
0
10
124
10
转发到社区