TL;DR: A new pipeline automatically builds 5,545 RL training tasks for coding agents using only source code itself, no issues or commit history needed, and it prioritizes quality over quantity.
Title: CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
URL:
Key points
🏗️ Auto-builds 5,545 tasks from 3,185 repos across 23 languages and 15 domains
🧪 Generates verifiers via execution-grounded tests, running the reference solution to record expected outputs
🛡️ Three-stage filtering: leakage checks, agent-solution agreement, and rollout difficulty filtering
📈 Big gains after RL training: DeepSWE +11.7pt, ProgramBench +17.0pt, Terminal-Bench +8.5pt
🔍 5k filtered tasks consistently beat 8k unfiltered tasks, proving quality beats quantity
🧠 Trained agents explore more and self-verify more, and these behaviors transfer to external benchmarks
I like how simple and practical the core idea is: you don't need issue trackers or dev history to build RL environments, just the code.
#
CodingAgent# #
ReinforcementLearning#