TL;DR: A new pipeline automatically builds 5,545 RL training tasks for coding agents using only source code itself, no issues or commit history needed, and it prioritizes quality over quantity.
Title: CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
URL:
Key points
๐๏ธ Auto-builds 5,545 tasks from 3,185 repos across 23 languages and 15 domains
๐งช Generates verifiers via execution-grounded tests, running the reference solution to record expected outputs
๐ก๏ธ Three-stage filtering: leakage checks, agent-solution agreement, and rollout difficulty filtering
๐ Big gains after RL training: DeepSWE +11.7pt, ProgramBench +17.0pt, Terminal-Bench +8.5pt
๐ 5k filtered tasks consistently beat 8k unfiltered tasks, proving quality beats quantity
๐ง Trained agents explore more and self-verify more, and these behaviors transfer to external benchmarks
I like how simple and practical the core idea is: you don't need issue trackers or dev history to build RL environments, just the code.
#
CodingAgent# #
ReinforcementLearning#