Register and share your invite link to earn from video plays and referrals.

Modular
@Modular
Building AI’s unified compute layer. We are hiring → 🚀
2 Following    24.3K Followers
Monday's community meeting showcases several standout community projects, including: • floki: an HTTP client for Mojo that works like Python's requests library, supported by the Modular Community Grant Program • noeira: an end-to-end physical AI stack in Mojo - physics, learning, perception and deployment, from simulator to robot, supported by the Modular Community Grant Program • warp: a look at building an async runtime to manage GPU synchronization calls across coroutines in Mojo Join us via Zoom at 10 AM PT:
Show more
This week in Chicago, Mojo found The Bean, ate a pretzel half its size, and got to hang out with the whole Modular team at our off-site. Want to come to the next one? We're hiring across Engineering, Product Management, Customer Engineering, and Developer Relations:
Show more
.@clattner_llvm takes the #SnapdragonSummit# stage for the first time as EVP Advanced AI Software @Qualcomm following the @Modular acquisition. Introducing the audience to the stack and it's role in the broader ecosystem.
Show more
GPU architecture | LLM Inference Handbook Add this to your LLM learning resource bundle. "Before writing or tuning GPU kernels, you need a working model of how a GPU runs code. Without it, suggestions like “increase occupancy” or "reduce shared memory bank conflicts" are just a set of rules to memorize. You don't fully understand when they apply and when they don't. This section explains modern GPU architecture at the level needed for kernel work. The details lean toward NVIDIA hardware because CUDA dominates much of the LLM inference ecosystem today. However, the core concepts apply broadly to AMD GPUs and other parallel accelerators as well."
Show more
Optimizing large scale inference systems is what we do, so we decided to write down what we know. Our LLM Inference Handbook is a free reference covering TTFT, TPOT, goodput, continuous batching, chunked prefill, prefix caching, KV cache math, prefill-decode disaggregation, quantization, and more. It includes 20+ interactive visualizations, is updated continuously, and is open to PRs.
Show more
There’s no single answer for where AI goes next. At @modular’s ModCon, GV’s @davemuni joined leaders across AI and venture to talk about the next wave of opportunity. Dave’s view: the infrastructure is taking shape. Now, the exciting part is what founders build on top of it.
Show more
Looking to get started with MAX? At ModCon 2026, Ehsan M. Kermani @ehsanmok and Bingfeng Xia went layer by layer through the stack: MAX Serve, the framework, and the Mojo kernel library, ending with a live agent bringing up a new model end to end. Start here:
Show more
At ModCon this year, we asked five investors where the next wave of AI infrastructure capital is going: training, inference, or silicon? Five different answers, and one panelist said it's the wrong question to be asking. The discussion also covered whether chipmakers absorbing AI software means the infrastructure is maturing or getting ahead of itself, and what open weight models do to the valuation of a compute-heavy startup. Michelle Gonzalez (@M12vc), Liz Stein (@USITfund), Sam Fort (@dfjgrowth), Quentin Clark (@generalcatalyst) and Dave Munichiello (@GVteam) each closed with what they think founders should build right now. Full recording:
Show more
At #ModCon2026#, Qualcomm CEO @Cristianoamon shared a bold vision for the future of AI infrastructure—and why our acquisition of @Modular is a massive win for the global developer ecosystem.
Fragmentation creates a tax across the entire AI ecosystem. At @PyTorch Conference this year, @clattner_llvm presents an alternative: one software stack to unite heterogeneous hardware. Join us in San Jose on October 20-21: #PyTorchCon#
Show more
New chips are shipping faster than ever before, but the ecosystem is being held back by having to rewrite and then re-debug the same code over & over again. @clattner_llvm CEO and Co-founder at @Modular and EVP of Advanced AI Software & Platforms at @Qualcomm, will deliver a keynote at PyTorch Conference North America about an open software platform for heterogeneous compute powered by Mojo and MAX. PyTorch has always been the place where the best models come together, and now there's a way to get those models onto all kinds of hardware. If you're interested in Al and compute, join us at the PyTorch Conference in San Jose, CA. Register now: #PyTorchCon#
Show more
Headed to Santa Clara today for #AIInfraSummit#? Don’t miss @alisterburt's talk at 10:30 AM PT in Expo Theater 2: "Modular: Open Source, Open Cloud, Open Silicon." Stop by the Qualcomm booth (#206#) anytime this week during expo hall hours to chat with the Modular team and catch a Modular Cloud demo.
Show more
If you looked at Mojo in 2024, liked it, and decided to check back later, later is now. Last month, the language hit 1.0, beginning a new epoch of stability for Mojo. The full Mojo language also went open source under Apache 2.0, including the compiler and more. Now is a great time to start building with Mojo. Clone the repo, build the compiler yourself, and learn from its unique MLIR internals and full commit history. At ModCon 2026, Brad Larson (Staff PM, Mojo) and Denis Gurchenkov (Senior Director, Mojo Compiler Engineering) walked through why Modular built a language at all, what stability means in the 1.x series, and a hint of what's next:
Show more
Today's AI software is fragmented. Every accelerator has its own compiler, kernel libraries, runtime, and development path. Developers absorb the cost of this fragmentation. In his ModCon 2026 tech talk, Abdul Dakkak, Chief Scientist at Modular, presents our alternative: a unified compute layer, flexible enough to extend to new models, modalities, and hardware. Abdul walks through our recipe for bringing up new hardware, demonstrated with three examples: 1. AWS Trainium, via our own team. Gemma 4 31B end to end. 2. Google TPU v6e, through our partner @HTECgroup. Their team brought it up with no LLVM backend to lower to, no prior knowledge of Mojo or MAX internals, and minimal support from us. 3. d-Matrix Corsair, implemented by @dMatrix_AI's own team on their own stack. Repo access Tuesday, working matmul by the following Monday. Thanks to HTEC and d-Matrix for their collaboration, and to Mihailo for presenting HTEC’s learnings. Read about the HTEC collaboration: Watch the full talk:
Show more
Modular's price-performance on @Zai_org's GLM-5.2 (Non-reasoning) lands right on the Pareto frontier in @ArtificialAnlys' latest benchmark: near-top speed without the near-top price tag. We're just getting started, and we're ready for GLM-5.3. Expect to see a lot more incredible results. 🚀
Show more
We're the #1# trending repo on @github today. The Mojo compiler went open source at ModCon this week, and developers noticed fast. Thank you to everyone who joined us in person and online. If you missed the keynote, the full recording is live on our YouTube channel:
Show more
Every time a new AI accelerator enters the market, developers face the same question: how much of our stack do we have to rewrite? Modular’s unified compute platform is making that answer clear: none. Qualcomm Technologies' data center AI accelerators are now integrated into the Modular Platform. Blog:
Show more
Four and a half years ago we made a bet: AI wouldn't run on one kind of silicon forever, and the software stack would have to be rebuilt for that world. Today at #ModCon2026#, we showed what that bet has become. Our biggest announcements from the keynote in the thread. 👇
Show more
Today, we open sourced Mojo 🔥. Announced just now during the ModCon keynote, effective immediately, Apache 2.0 License. Thank you to our community for waiting patiently and building alongside us. #ModCon2026# Full blog:
Show more
0
46
1.4K
225
Forward to community
Our biggest product announcements of the year land tomorrow in the ModCon 2026 opening keynote. 9:00 AM PT. Chris Lattner, Tim Davis, Eric Johnson, and Mostafa Hagog deliver The Unified AI Compute Layer, alongside Cristiano Amon and Rashid Attar of @Qualcomm, Anush Elangovan of @AMD, and Darko Todorovic of @HTECgroup. If you only watch one thing from ModCon, watch this. Free livestream:
Show more