Register and share your invite link to earn from video plays and referrals.

Search results for OML
OML community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including OML
oMLX 0.7.0rc1 is out! This release brings faster Qwen prefill & generation, MiMo V2.6, Ternary Bonsai 2, and partial block caching. DFlash now handles concurrent requests together, and Lightning MTP gets faster batch decoding! Performance on M5 Max, 128 GB (Prefill, oQ4e quant) - Qwen3.8-Flash-Next: 1,522 -> 2,007 tok/s (+32%) at 16K context. (Decode, batch=4, oQ4e quant) - Qwen3.8-27B with DFlash2: 56.9 -> 131.5 tok/s (+131%) - Qwen3.8-27B with Lightning MTP: 88.9 -> 136.9 tok/s (+54%) Full benchmark details are in the release notes. New models and features - MCDMA RDMA support for Mac + CUDA deployments, contributed by @ashxhart. - Partial block caching. No more reprocessing thousands of tokens just because they didn't fill a complete cache block. In one test, next-turn prefill dropped from 1,174 tokens to 37. - Ternary Bonsai 2 text and vision support. - MiMo V2.6 image, video, and audio understanding, plus Lightning MTP and DFlash for compatible checkpoints. - Broader MoE expert offload, including Lightning MTP alongside expert offload for DeepSeek V4.1 and GLM-5.3-Flash. This RC also includes the improvements from the dev releases, including Cluster v2, one-click model settings from community benchmarks, and a customizable dashboard. The GDN prefill kernels are adapted from @ddalcu's excellent mlx-serve! After a short round of testing, I'll publish the stable release and keep moving forward!
Show more
oMLX 0.7.0.dev4 adds DeepSeek V4.1 CED prefill with up to 79% faster prompt processing on M3 Ultra (opt-in), multi-request Lightning MTP, and plenty of quality-of-life improvements including one-click model settings. You can now apply model settings with one click using over 450,000 community benchmarks on Pick a top result for your Mac chip and model to use the same settings. Using a customized model? Copy a one-line recipe from a model with the same architecture and paste it into oMLX. There's also a customizable dashboard, global settings reset, and many bug fixes. See the release notes for the full details. Thank you for contributing code and sharing your benchmarks. Your benchmarks now help other users find and apply settings for their models with one click! * This release upgrades mlx-lm and mlx-vlm, with extensive internal changes. If something that worked before breaks, please open a GitHub issue with logs, and roll back to dev2 for now. ** I test as much as I can before each release, but I can't cover every Mac configuration and model. Issue reports on dev builds are a huge help to me.
Show more
oMLX community benchmark is back up, home to the largest collection of Apple Silicon MLX benchmark results! I got way more submissions than I expected, and my old plan couldn't keep up. So I improved the search logic and upgraded the plan. Depending on how traffic goes I may need to tune it more, but for now I'm just praying it holds 🙏 I'm also thinking about ways oMLX and the community benchmark can work better together. Stay tuned!
Show more
oMLX hits 47 tokens per second on a base M2 MacBook Pro by offloading context to the SSD. We explore how native MLX features achieve 3x faster generation than LM Studio in our latest test.
Show more
OML LB got me like: One wrong spam and poof - deranked to rat tourist status . Pro move: Thoughtful engages + original shitposts on BTC hustles. Gated Rune via RadFi, 1B supply - join before the streets overflow. @__O_M_L__ #OML#
Show more