A chunk of the RL community really doesn’t like replay buffers. Concerns about the cost of storing a million observations are valid, but the recent evangelization of “streaming” RL, where observations are just used a single time and discarded, feels like a poor design point to me.
It is interesting that it can now work at all, but I’m pretty sure that the optimal number of saved observation buffers is not “one” at any memory constraint. There would certainly be useful things to do with a managed buffer of just hundreds of sparse observations, even if you didn’t do bootstrapping from them.