๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Guanlan Dai
@guanlan
Building @Runta, the execution layer that controls what AI agents can actually do. Prev @Cloudflare and @Kong.
๊ฐ€์ž… January 2010
478 ํŒ”๋กœ์ž‰ ์ค‘    6.2K ํŒฌ
A year ago the question was which model. Now it's which harness. Pi, Exo, Claude Code, Codex, DeepSeek Harness and 4 others. Same model, same tasks, same runtime. 360 runs, 2 billion tokens. Pass rates: 50% to 67%. Cost per pass: $1.05 to $18.34. Introducing FrontierHarness Eval. ๐Ÿงต
๋” ๋ณด๊ธฐ