注册并分享邀请链接,可获得视频播放与邀请奖励。

justin
@justinsunyt
cofounder @capydotai • prev @pennmandt
加入 May 2023
1.7K 正在关注    8.8K 粉丝
actually, all three beliefs are right the less harness, the better: true for frontier models working on isolated single shot tasks. mini-swe-agent outperforms codex and claude on coding benchmarks with a single bash tool, while being significantly cheaper. for tasks that don’t interact with external systems, models benefit more from less context bloat than more tools and instructions. post training wins: true for specific tools and workflows. the RL-ification of models has led them to overfit on tool schemas. hence why opus hallucinates parameters on the pi edit tool, or gpt will prefer native apply_patch. and as the labs invent new approaches to multi-agent orchestration, model behaviors around spawning subagents will branch out too. harnesses provide more value: true if you care about agent behavior around interfaces and systems. as agents evolve and move towards the cloud, we will see infinitely more permutations of agent interfaces and tools. your agent in slack shouldn’t behave the same way as it does in your terminal. some of these behaviors aren’t as easy to tack on as an AGENTS.md or a skill. even more, the labs’ incentives aren’t always aligned with the end user. they want to sell more of their tokens. but the pareto frontier of models today is spread across multiple model families. the best harnesses will take advantage of model capabilities at each price point and route automatically. these three beliefs don’t have to conflict with each other at all. the ideal harness is minimal when it needs to be (deferred context), leans into model post training (model-specific guidance/tools), and brings more value by routing the best models and aligning to their interfaces. (running evals on capy v2 right now and seeing harness optimizations pay off in real time)
显示更多
0
10
97
10
转发到社区