Today's AI software is fragmented. Every accelerator has its own compiler, kernel libraries, runtime, and development path. Developers absorb the cost of this fragmentation.
In his ModCon 2026 tech talk, Abdul Dakkak, Chief Scientist at Modular, presents our alternative: a unified compute layer, flexible enough to extend to new models, modalities, and hardware.
Abdul walks through our recipe for bringing up new hardware, demonstrated with three examples:
1. AWS Trainium, via our own team. Gemma 4 31B end to end.
2. Google TPU v6e, through our partner
@HTECgroup. Their team brought it up with no LLVM backend to lower to, no prior knowledge of Mojo or MAX internals, and minimal support from us.
3. d-Matrix Corsair, implemented by
@dMatrix_AI's own team on their own stack. Repo access Tuesday, working matmul by the following Monday.
Thanks to HTEC and d-Matrix for their collaboration, and to Mihailo for presenting HTEC’s learnings.
Read about the HTEC collaboration:
Watch the full talk: