Our COLT 2026 paper!
A single stepsize with high probability gives both fast and robust rates in TD learning, without using any projection.
IMO, Wei-Cheng put together a very interesting proof, with some new tricks as well.
1/11
New paper: A Single Stepsize Suffices for Unprojected Linear TD(0)
Can TD(0) without projection and knowledge of curvature achieve high-probability rates that are both robust and fast under Markovian sampling?
We show: one curvature-free stepsize + PR average suffice. 🧵