Register and share your invite link to earn from video plays and referrals.

Prince Canuma
@Prince_Canuma
Creator of (@Nativ_AI, mlx-audio & mlx-vlm) • working on something new Ex-@arcee_ai@neptune_ai
1.3K Following    22.8K Followers
highlight of my Warsawmaxxing 🇵🇱 was discussing local AI with @Prince_Canuma 🫡
Had the pleasure of meeting and chatting with @tugot17 at great length about local ai, hardware architecture, his latest release DSPark for LFM models and more! One of the smartest people I’ve met in Poland 🇵🇱
Show more
Congrats to @Zai_org on GLM-5.3-Flash! Day 0 support in Nativ 🎉 320B params, 18B active · 1M native context · fully on-device. On an M3 Ultra (512GB) with Nativ v0.3.4, standard 4-bit MLX: ⚡ Up to 505 tok/s prefill · 32 tok/s decode 🚀 Batch 4 lifts decode to 70.9 tok/s 💾 Peak under 380GB, fully resident, no offload 📏 Benchmarked to 128K context A 320B MoE running entirely on your Mac. Try Nativ 👇
Show more
.@Alibaba_Qwen announces Qwen3.8-Flash-Next (Qwen 4 architecture), a new open-weight multimodal MoE model. The model will be released today and will come with Day-0 support on @Nativ_AI
The next-gen architecture powering Qwen4 is now here! ✨ Get ready for the open release of Qwen3.8-Flash-Next 🚀 The countdown starts now! ⏳🔥
I would love to get my hands on a M5 Ultra 🙌🏽
Great day for Local AI, M5 Ultra and M6 Macs are out! 🔥 M6: dual 16-core Neural Engines + 170GB/s bandwidth. 
M5 Ultra: up to 80 GPU cores, 512GB unified memory and 1.2TB/s bandwidth—enough to run massive LLMs entirely on-device. The AI workstation is becoming a Mac.
Show more
Great day for Local AI, M5 Ultra and M6 Macs are out! 🔥 M6: dual 16-core Neural Engines + 170GB/s bandwidth. 
M5 Ultra: up to 80 GPU cores, 512GB unified memory and 1.2TB/s bandwidth—enough to run massive LLMs entirely on-device. The AI workstation is becoming a Mac.
Show more
Coming to MLX-VLM and @Nativ_AI 🚀
Meet Apodex 1.1: Scaling Agentic Intelligence for Complex Work Open Source Harness: Open Weights: We’re excited to introduce Apodex 1.1, our new model family built to scale agentic intelligence for professional work. 🧠 Frontier-level intelligence for complex work Apodex 1.1 brings frontier-level agentic performance across complex professional work, scientific research, financial analysis, and deep search. 🤝 Asynchronous Agent Team Apodex 1.1 can break down complex tasks, coordinate multiple agents in parallel, continuously integrate their findings, and let you step in to guide or redirect the work at any time. 🔬 Open-source research workbench We’re open-sourcing FrontierAgent—a locally deployable research workbench for the Apodex 1.1 family, including asynchronous Agent Team. Available now: 🔹 Apodex 1.1 — our most capable frontier model, available through the Apodex online workbench 🔹 Apodex 1.1 mini — open-weight model for running complex work locally 🔹 FrontierAgent — open-source, locally deployable research workbench 🌐 Live workbench: 🦾 API platform: 📃 Paper:
Show more
Faster by the day @Nativ_AI 🚀
Nativ v0.3.4 is here 🚀 ⚡ Faster startup and smoother app 🎙️ Audio-file imports — drop recordings into the Audio Library and get automatic transcription + summaries. ⚙️ Live server settings — inspect and change engine settings from the Developer page, no restart needed 📊 Richer system stats — temperature, fan speed, thermal pressure, and power usage in chat 🎨 Control-panel redesign — smoother sidebar resizing, consistent materials, better fullscreen 🔌 Better MCP diagnostics — failed connections show clear, redacted errors instead of hanging on "Connecting" 🔐 Hugging Face tokens now live in macOS Keychain Plus a few UX fixes: gated-model errors, multi-key dictation shortcuts, and clearer failure logging. Download Nativ now:
Show more
Nativ v0.3.4 is here 🚀 ⚡ Faster startup and smoother app 🎙️ Audio-file imports — drop recordings into the Audio Library and get automatic transcription + summaries. ⚙️ Live server settings — inspect and change engine settings from the Developer page, no restart needed 📊 Richer system stats — temperature, fan speed, thermal pressure, and power usage in chat 🎨 Control-panel redesign — smoother sidebar resizing, consistent materials, better fullscreen 🔌 Better MCP diagnostics — failed connections show clear, redacted errors instead of hanging on "Connecting" 🔐 Hugging Face tokens now live in macOS Keychain Plus a few UX fixes: gated-model errors, multi-key dictation shortcuts, and clearer failure logging. Download Nativ now:
Show more
This is fast! LFM2.5 is a great model series and now it’s even faster!
LFM2.5 DSpark by @liquidai is coming to mlx-vlm in v0.6.16 ⚡️ Exact speculative decoding on M5 Max, delivering up to 3.7× faster generation with zero output drift. On-device speed, zero output drift. Benchmarks below 🧵
Show more
Can’t wait for OSS drop 🔥👀
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
Show more
Got @Nativ_AI installed my M1 MacBook Pro with 16GB of RAM. Poking around and seeing how far I can get with it on my MacBook Pro. Next installation will be my 1st gen Mac Studio which has an M1 Pro with 32GB of RAM. Thanks to @RayFernando1337 and @Prince_Canuma for such a really good Local AI Masterclass livestream. #LocalAI# #NativAI#
Show more
Some initial results with @liquidai LFM DSpark
Coming to MLX-VLM and soon Nativ 🚀
Coming to MLX-VLM and soon Nativ 🚀
Used @Nativ_AI voice dictation (powered by @cohere transcribe) and instructed FLUX.2 by @bfl_ai to generate images on my Mac 💻 All fully offline, 35,000 ft up, on my way to my home country Mozambique 🇲🇿
Show more
Congrats to @liquidai on LFM2.5-2.6B! Excited to have partnered with them for Day 0 support in Nativ 🎉 Built for agentic + coding workflows — and it’s fast. On an M5 Max (48GB) with Nativ v0.2.2 — full bf16, no quantization: ⚡ 11,231 tok/s prefill ⚡ ~84 tok/s decode 🧠 Full 128K context in just 8.5GB 📈 476 tok/s aggregate decode at batch 16 Local inference doesn't get much better than this. Get started today👇🏽
Show more