๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Yuhang Yao
@yuhang_yao
Senior Research Scientist @ZOOM | PhD @CarnegieMellon | Graph Agent Research
๊ฐ€์ž… December 2021
117 ํŒ”๋กœ์ž‰ ์ค‘    1K ํŒฌ
GLM-5.2 is now just 0.005 away from Fable-5 in visual spec-to-app development. On VISTAโ€™s latest C4 leaderboard: ๐Ÿฅ‡ Claude Code + Fable-5 โ€” 0.274 ๐Ÿฅˆ Claude Code + GLM-5.2 โ€” 0.269 ๐Ÿฅ‰ Claude Code + Opus 4.8 โ€” 0.263 ใ€€ Claude Code + Sonnet 4.6 โ€” 0.248 ใ€€ Cursor + Composer 2.5 โ€” 0.212 ใ€€ Codex + GPT-5.5 โ€” 0.205 The story is no longer just โ€œwhich model writes better code.โ€ For real app-building agents, performance is increasingly shaped by the full stack: ๐Ÿง  model ๐Ÿ› ๏ธ harness ๐Ÿ” tool loop ๐Ÿ‘๏ธ visual understanding โš™๏ธ reasoning setup The harness is becoming part of the model. ๐Ÿ”—
๋” ๋ณด๊ธฐ