註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Scott Condron
@_ScottCondron
Helping build AI dev tools at @weights_biases. I post about AI, data visualisation and the stuff I’m working on at wandb.
加入 April 2018
2.1K 正在關注    5.8K 粉絲
TLDR all harnesses are basically the same because they all need to preserve the sacred KV cache hit rate of the proprietary models BUT what if you control the underlying infrastructure and use open models? Then, instead of just using whatever harness you’re given, your harness and inference scheduler are aware of each other so… There’s lots of infra optimizations that can be done :) Session affinity is the obvious one but there’s a lot more - speculative tool calling - request priorities based on workload - prefetch KV cache during tool calls - colocate subagents and their tool calls - prewarm tool calling sandboxes the list goes on! I’m no expert here but know some folks that are (lmk if you’re interested in learning more)
顯示更多
0
27
639
31
轉發到社區