註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Yuhang Yao
@yuhang_yao
Senior Research Scientist @ZOOM | PhD @CarnegieMellon | Graph Agent Research
加入 December 2021
117 正在關注    1K 粉絲
GLM-5.2 is now just 0.005 away from Fable-5 in visual spec-to-app development. On VISTA’s latest C4 leaderboard: 🥇 Claude Code + Fable-5 — 0.274 🥈 Claude Code + GLM-5.2 — 0.269 🥉 Claude Code + Opus 4.8 — 0.263   Claude Code + Sonnet 4.6 — 0.248   Cursor + Composer 2.5 — 0.212   Codex + GPT-5.5 — 0.205 The story is no longer just “which model writes better code.” For real app-building agents, performance is increasingly shaped by the full stack: 🧠 model 🛠️ harness 🔁 tool loop 👁️ visual understanding ⚙️ reasoning setup The harness is becoming part of the model. 🔗
顯示更多