๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
280 ํŒ”๋กœ์ž‰ ์ค‘    415 ํŒฌ
TL;DR: Three failure modes of long-horizon agents โ€” compounding errors, context rot, and task-state loss โ€” solved structurally through a Manage-Execute-Audit (MEA) loop. WeaveBench PassRate: 51.8% โ†’ 80.7%. LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Key points: ๐Ÿ”ง Manager: Maintains explicit task state outside the execution trajectory. Constructs subtask contracts specifying goals, acceptance criteria, constraints, and evidence. โšก Executor: Runs each subtask in a fresh, budget-bounded context โ€” no prior trajectory history passed in. ๐Ÿ” Auditor: Post-execution read-only inspection, independent of the executor. Reports completion, integrity, and state updates; serves as persistent cross-round memory. ๐Ÿ–ฅ๏ธ GUI/CLI hybrid: Manager routes tasks to the right interface; AgentAdapter swaps in Claude Code, Codex CLI, Hermes Agent without modifying native loops. ๐Ÿ“Š WeaveBench: PassRate 51.8% โ†’ 80.7%; Design +60pp, Spatial/3D +50pp. ๐Ÿค– OSWorld 2.0: Qwen 3.7-Plus 2.8% โ†’ 8.3% (3ร—); Claude Opus 4.7 20.6% โ†’ 35.3%. ๐Ÿ’ป Terminal-Bench: 69.7% โ†’ 77.2%, with 24% fewer tokens consumed. ๐Ÿ’ก Manager overhead: only 2โ€“8% of total tokens. Auditor carries 19โ€“38% โ€” the primary cost of reliability. "Agent capability is a property of the complete modelโ€“harness system, not the model alone" โ€” that framing changes how you should think about long-horizon agent design. #AIAgents# #LLM#
๋” ๋ณด๊ธฐ