Highlights from the
@aiDotEngineer keynote chat with
@olive_jy_song from
@MiniMax_AI, creators of the open-weight M3 model:
* An intern came up with sparse attention
* The minimax models have "native" multimodality (trained from the start)
* "Could we have models with 1 trillion context window?" "We could explore that..."
* M3 is building M3.1 right now