๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
280 ํŒ”๋กœ์ž‰ ์ค‘    415 ํŒฌ
๐Ÿงฉ What autonomous AI agents are missing isn't a smarter model โ€” it's on-the-ground know-how. That's the premise of this paper. Title: Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills URL: โ“ What's missing from autonomous research agents? ๐Ÿ’ก The usual two-layer view โ€” model plus execution harness โ€” leaves out the operational knowledge of picking the right method, using package APIs correctly, and avoiding implementation pitfalls. This paper treats that as an explicit third layer. โ“ How do you actually get that knowledge? ๐Ÿ’ก It distills GitHub repos and papers through a four-stage pipeline โ€” Scope, Ground, Construct, Verify โ€” into verified "skills." From 1,000 repos and 153 papers, they built a library of 5,353 skills. โ“ How much difference do skills actually make? ๐Ÿ’ก With the same GPT-5.5 backbone and same harness, just adding skills lifts MLE-bench from 31.11% to 72.89%, with similar gains across PaperBench, FrontierCS, and PassNet โ€” hard tasks see over 4x improvement. โ“ Isn't this just throwing more compute at the problem? ๐Ÿ’ก No โ€” improvement barely correlates with token counts or tool calls, and it beats a Claude Opus 4.8 setup while using fewer tokens. The knowledge itself is doing the work. #AIAgents# #MachineLearning#
๋” ๋ณด๊ธฐ