We also investigated different amounts of distillation and the rate of accuracy improvement that distillation confers, amongst other results.
This is a first step toward understanding how different forms of distillation affect RL, and how to warm-start model training more effectively.
Full work: