Register and share your invite link to earn from video plays and referrals.

Awni Hannun
@awnihannun
ow knee
358 Following    45K Followers
@jundotkim makes local inference awesome on Mac with oMLX, really glad to see Hugging Face accelerating the work. The community are building great projects with MLX, and they have been motivating us to keep improving it.
Show more
I'm happy to announce that I've joined Hugging Face. What started as a personal project back in February is now something I get to work on full time. Local AI has grown explosively this year. I've said this since the early oMLX releases: I want my friend who bought a MacBook yesterday to be able to run AI on it today. I believe MLX has that potential. Apple Silicon is the easiest entry point for a regular person to get started with AI, with no complicated hardware to assemble. And the open source community, including Hugging Face, has been growing that potential. That support has already made a huge difference. Hugging Face is the best place for me to support oMLX and the MLX community with everything I have. For anyone getting into local AI, the first step usually starts at Hugging Face. Mine did too. I'm proud that I get to work at that entry point, where new ideas can be tried out. oMLX stays exactly where it is, under the same Apache 2.0 license in the same repository, and I'll keep leading the project, same as before. What changes is that I can now spend far more time on it, move faster, and build something sustainable for the long term together with all the contributors who have put so much into it. I also want to do more for MLX as a whole. oMLX is built on top of transformers, mlx-lm, mlx-vlm and the rest of the ecosystem, and I'm grateful to the people behind them. Rather than keeping everything inside oMLX, over time I want to push work upstream where it makes sense. And where the community needs something that doesn't exist yet, oMLX is a good place to try it first. Thank you to the 264 contributors who have built oMLX with me, and to everyone who filed issues with detailed logs and reproductions so we could fix things. oMLX would not be what it is without you. oMLX continues in the same place, in the same way. Just much faster. HF's announcement:
Show more
0
130
1.1K
101
Forward to community
Super happy to welcome @jundotkim, oMLX creator, to Hugging Face! 🥳 We are invested in Local AI and support MLX since @awnihannun and @angeloskath released it in Christmas 2023. oMLX gets stability and speed. It remains a community project led by Jun.
Show more
> Apple already offers a tool for efficiently running AI models on Apple hardware called MLX Let's go!
$47.9 million mistake not naming the company 🫶 Heart Hands
JUST IN: NVIDIA's $12,930,300,000 acquisition of Hugging Face contains an easter egg. The number 129,303 is the decimal conversion of Unicode point U+1F917. The 🤗 emoji.
Today we are releasing our speculative decoding implementation in Uzu. Biggest release since inception of our lab. Initially, for Qwen3.6 27B, with support for Qwen3.8 27B and Muse Glimmer coming soon.
Show more
Qwen 27B dense at 105 tok/s output on an m5 max is pretty bonkers. Breaking down the memory wall one brick at a time.
We’re releasing speculative decoding in our inference engine Uzu, starting with Qwen3.6-27B. On Apple M5 Max with 128 GB of unified memory, our Mirai-M stack reaches 105 output tokens/sec entirely on-device - 2.9× faster than the fastest MLX speculative-decoding implementation we benchmarked. The result comes from full-stack co-design: DFlash + our Weaver model, tree-based speculative decoding, Mirai quantization, our verification algorithm, and Metal kernels optimized for Apple silicon. Mirai builds the full stack for frontier on-device intelligence. 105 t/s is the evidence, full-stack integration is the moat. But tokens/sec are not actually the metric we ultimately care about. More on that soon. Full benchmarks + methodology:
Show more
Starting today, 10,000 scientists across every field, from math to chemistry to physics and more, can get Claude through our new Claude Team plan for scientists. Standard seats are free, and premium seats with 5x usage limits are $15 per month, an 80% discount, for one year. Claude is becoming increasingly capable of scientific work, with recent progress on problems from advanced physics calculations to protein design. Alongside that progress, we've been investing in the research community: Claude Science launched in June, and our AI for Science program funds high-impact projects with free credits. Today's expansion builds on both. Principal investigators (or equivalent) at academic and nonprofit research institutions can sign up, then add the researchers in their group. Over the coming months, we plan to extend the program well beyond the initial 10,000 seats. Learn more:
Show more
0
540
12.6K
1.2K
Forward to community
exo featured on Apple's new M5 Ultra Mac Studio and M6 / M5 Pro Mac Mini pages. Over the past year, we have worked closely with Apple on low-latency RDMA networking over Thunderbolt 5, enabling clusters of Macs to run massive models like Kimi K3 and GLM-5.3 at API speeds. With RDMA, aggregate memory bandwidth across Macs scales ~linearly. A cluster of 4 x M5 Ultra Mac Studios scale to an aggregate memory bandwidth of ~4.8TB/s. Previously, these were speeds only achievable with data center GPUs. Our vision is a data center on every desk. Apple Silicon's superior memory unit economics, power efficiency, and out of the box experience for Local AI make that possible. Thank you to @angeloskath, @awnihannun, @twid and countless others at Apple who tirelessly to bring this technology to the world.
Show more
0
63
1.4K
117
Forward to community
MLX v0.32.2 is here. It has only been a week since last release, but there are some nice optimizations benefiting both decoding and prefilling that we would like you to enjoy early:
Show more
Thoughts on mlx-lm: top-priority is making it the central registry of MLX model implementations, with tools for evaluation and profiling, and we should add vision models too. Inference engines can have their own schedulers and custom kernels, and do whatever hack to make inference ultra fast, while importing mlx-lm as a library of models. We can rely on community to contribute model implementations, but there would be a fixed procedure to verify correctness of the implementation, ideally automatically. Everything else except for critical bugs, should be irrelevant at the moment and I'm closing PRs and issues aggressively. Many people will be mad at this, and certainly I would be making mistakes closing legitimate things, but for the project to survive, and for the community to grow healthy, I don't see another way.
Show more
MLX v0.32.1 has been released, with a lot of minor fixes and performance improvements thanks to the contributors. Last 2 releases also received much more external contributions than ever.
Show more
@ronaldmannak Ollama’s latest inference on Apple Silicon is also built on MLX: MLX is transformative.
Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far. Fastest right now: Ollama's MLX engine (DFlash included) at ~29 tok/s. Tuned llama.cpp: ~21. Raw mlx-vlm: ~10, not optimized yet. Numbers below if you're setting it up 👇
Show more
I was working on some optimization for mlx and noticed that our rmsnorm backward was hitting significantly lower bandwidth than forward. I wrote a small note on why that’s the case and how it could be fixed.
Show more
Muse Glimmer is now available to run with Ollama. Available today via Ollama’s MLX engine with state-of-the-art-performance on Apple Silicon, Muse Glimmer can power Claude Code, Codex, and more always-on local agent workflows natively using Ollama. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other platforms will be available shortly. Go @finkd @alexandr_wang and @Meta !
Show more
0
73
1.3K
108
Forward to community
Run Kimi K3 on a Mac Studio 🫰 K3 is 2.8T parameters and 1.6TB on disk, which makes it impossible to run on Apple Silicon. Until now. Our MLX port is now open source: To accomplish this, we solved two things: 1. We wrote a streaming converter that walks one layer at a time, so that mlx_lm doesn't need to materialize the whole model. 2. REAP pruning sits on top and scores all 896 experts against a calibration corpus to keep only ones your workload needs. That's what brings K3 down to 350GB and inside a Mac Studio.
Show more
0
123
3K
319
Forward to community
Excited to launch with @eigenlabs today. It's an open autoresearch competition to make Laguna XS 2.1 inference as fast as humanly (and agentically) possible on consumer Macs. Eigen's agents already found 36.8% faster inference, and that's before the competition even started. The best part of open weights is that the community takes a model further than any of us could. 
Can't wait to see what everyone does on the leaderboard! 
Show more
🚀 Rapid-MLX 0.11.0 is out! Making local models on Apple Silicon reliable enough to run your agent workflows, not just demo them. ⚡️ Performance • Prefix-cache reuse: 13.1s → 0.51s TTFT (25.6x faster on a repeated 6k-token prefix). Shared system prompts stop paying prefill twice. One session saved 38,032 tokens. • Response cache: 656ms → 2ms (284x faster, byte-identical, zero GPU decode). • Throughput: 152 tok/s on Qwen3.5-4B-4bit. Cold starts in seconds. 🤖 "rapid-mlx chat" grew into a real agent Point it at any MCP server and it chains tools autonomously (e.g., filesystem server → 14 tools → list_directory → answer), complete with a live tool-activity UI. Thinking budgets force-close at decode time to keep things snappy. 🧠 5 New Model Families • HY3 295B MoE (native MTP speculative decoding), Qwen3-Coder-Next 80B, MiniCPM5, LFM2.x, and a 2-bit Ternary Bonsai that actually runs. • Kokoro TTS → Parakeet STT round-trip verified word-perfect. • Qwen3-VL vision, KV cache export/import, and structured output on llguidance. 🛠️ Schema-valid tool calls by construction. If you run coding agents on local models, you know the pain of "almost-JSON" arguments killing your parser. 0.11.0 constrains decoding itself—the model cannot emit a malformed tool call. • Verified across harmony/gpt-oss, qwen3_coder_xml, gemma4, deepseek_v3, plus hermes. • Auto, forced, and required. On by default, zero flags. • The Keystone: It now holds for reasoning models too! tags no longer break forced calls because of grammar offset issues. 🛡️ Reliability & Honest Gaps Every release runs an automated dogfood battery (5 models × 11 checks on real hardware). Invalid schema → honest 400. Unknown model → 404 with available options. No silent fallbacks. Gap: Qwen3.6 native MTP is parked for now, and a same-machine Ollama head-to-head is up next. ⬇️ Now in homebrew-core brew install rapid-mlx rapid-mlx chat qwen3.5-4b-4bit Apache-2.0, MLX-native. 3.4k stars — break it and tell us, issues get answered fast. ⭐️
Show more
Excited to introduce Nativ 🚀 Run frontier open models locally on your Mac. No accounts, no subscriptions, no cloud. ⚡ Built on mlx-vlm — multimodal + fastest on Apple Silicon 🔒 100% private — every token generated on your machine 🛠 Plug your coding agents into a local endpoint 📊 Live telemetry — tokens/sec, memory, usage 🆓 100% open source, MIT licensed, free forever Your Mac is more capable than you think. Stop renting intelligence. Download for macOS 👇 Github repo 🌟
Show more
0
108
973
140
Forward to community