My 2 cents:
Sparse attention probably makes continual learning a greater necessity. I suspect we could have achieved AGI/ASI easier with dense+RLMs, but at the cost of unacceptably unwieldy data generation. So Wenfeng's roadmap is perhaps the only feasible one.