๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Alex Ker ๐Ÿ”ญ
@thealexker
code+words @baseten | investing in frontiers & sharing my curiosities | prev @bloombergbeta @stanfordhai @neurable.
๊ฐ€์ž… May 2018
1.3K ํŒ”๋กœ์ž‰ ์ค‘    13.3K ํŒฌ
recent notes/learnings on harnesses (small n, directionally correct, experimental results soon): > running harness experiments is brutally expensive! e.g. @guanlan's experiment used 2B tokens but only covered 30 tasks (21 terminal-bench tasks and 9 deepswe tasks) > the best harness is one that's the most bitter-lesson pilled and cost-efficient. you do this by letting your harness adapt at runtime (rsi) and align itself with actual tasks and data. common failure mode is being too opinionated a priori or over optimizing for one model. good properties of harnesses have been mostly commoditized/converged. > as horizons of execution get longer, two things stand out: 1) speed (does the harness know when to use subagents, fanouts, and have good primitives around them). 2) the ability to unblock itself. this means not getting stuck in a particular loop or deadzone but stepping outside without human intervention. what do you look for in good harness design?
๋” ๋ณด๊ธฐ