The mates are in agreement: we are looking at mainstream, beyond transformers architectures taking significant market share this year.
@alexwg highlights MoEs, diffusion transformers, linearized attention, and recurrence as a disassembling of "Attention Is All You Need" until nothing original is left. I would add that
@liquidai already proved there are many new, discoverable architectures. There are simply too many ways in which vanilla transformers are grossly inefficient.