What happens when you let AI agents autonomously run research on recommender models that take days to train? This paper finds out.
Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender Systems
❓ Why aren't existing autonomous research agents enough?
💡 Prior systems assume feedback loops of minutes to hours. Industry-scale recommender models take days per training run and involve thousands of lines of configuration, so serial iteration breaks down and you need a fundamentally different design that runs many ideas in parallel across distributed servers and recovers from failure.
❓ How does it juggle parallel experiments and recover from crashes?
💡 Each idea keeps its own independent state file moving through ideating, implementing, validating, training, and analyzing. When a session restarts, it rereads a global registry and trajectory logs on shared storage, letting work resume seamlessly across servers.
❓ Does the system actually get better over time?
💡 Experiment trajectories are distilled into model-specific playbooks that record dead ends as "do not" rules. Across 31 iterations, major fixes needed per iteration dropped from 4.0 to 0.5, with 83% of iterations completing with zero fixes.
❓ How autonomous is it in practice?
💡 One session executed 970 consecutive log entries with zero human intervention, and the system even switched its own training workflow after repeated publish failures.
#
AIAgents# #
MachineLearning#