๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
280 ํŒ”๋กœ์ž‰ ์ค‘    415 ํŒฌ
TL;DR: A new pipeline automatically builds 5,545 RL training tasks for coding agents using only source code itself, no issues or commit history needed, and it prioritizes quality over quantity. Title: CodeMidas: Scaling Agentic Coding RL Environments from Code Itself URL: Key points ๐Ÿ—๏ธ Auto-builds 5,545 tasks from 3,185 repos across 23 languages and 15 domains ๐Ÿงช Generates verifiers via execution-grounded tests, running the reference solution to record expected outputs ๐Ÿ›ก๏ธ Three-stage filtering: leakage checks, agent-solution agreement, and rollout difficulty filtering ๐Ÿ“ˆ Big gains after RL training: DeepSWE +11.7pt, ProgramBench +17.0pt, Terminal-Bench +8.5pt ๐Ÿ” 5k filtered tasks consistently beat 8k unfiltered tasks, proving quality beats quantity ๐Ÿง  Trained agents explore more and self-verify more, and these behaviors transfer to external benchmarks I like how simple and practical the core idea is: you don't need issue trackers or dev history to build RL environments, just the code. #CodingAgent# #ReinforcementLearning#
๋” ๋ณด๊ธฐ