Register and share your invite link to earn from video plays and referrals.

Anush Elangovan
@AnushElangovan
GPU SW @ AMD
221 Following    13.4K Followers
Here is a PR that adds a toast: Let the race to zero begin.
Halo Laptops are the fastest with Omarchy:
🙏 Thanks for the kind words. The power of OPEN is that you get to talk shop all day and geek out, and the incredible AMD ROCm team does the heavy lifting. 💪 + 🦾🦾
i think @AnushElangovan is really smart. he gets software and what's happening in ai development right now.
AI isn't a one-platform future. As models, frameworks, and infrastructure evolve, developers need more choice. AMD VP of AI Software @AnushElangovan and AI futurist Zack Kass discuss why open software, developer flexibility, and ecosystem choice will shape what's next.
Show more
🎉 Congrats to @Alibaba_Qwen on Qwen3.8-2.4T-A95B, one of the largest open-weight models released to date. 2.4T params, 95B active, 512 experts. Day-0 support in vLLM, verified on @NVIDIA and @AMD hardware. A ready-made 4-bit checkpoint per vendor, both out of the box: Inferact/Qwen3.8-2.4T-A95B-NVFP4, 1.32 TiB, one NVIDIA 8xB300 node Inferact/Qwen3.8-2.4T-A95B-MXFP4, 1.45 TiB, one AMD 8xMI355X node No conversion, no calibration on your side. Just vllm serve. Thanks to @Alibaba_Qwen for the weights and the collaboration, @NVIDIAAI and @AIatAMD for the joint kernel engineering, @inferact for the quantized checkpoints and vLLM integration, @digitalocean and @togethercompute for early testing, and the vLLM community. 🙌 🔗
Show more
FastFlowLM 1.0 Released Now As Part Of The @AMD @AIatAMD ROCm Umbrella Running LLMs on @AMDRyzen AI NPUs on Windows and Linux.
AI isn't a one-platform future. As models, frameworks, and infrastructure evolve, developers need more choice, not more lock-in. Watch AMD VP of AI Software @AnushElangovan and AI futurist Zack Kass discuss why open software, developer flexibility, and ecosystem choice will shape what's next. Watch the full fireside chat on 8/13.
Show more
The @AMD @AIatAMD ROCm Spur: Providing AI-Native, Rust-Based Job Scheduling
Real operators in AMD Strix Halo local AI: @Italianclownz — ROCmFPX formats/kernels + Hermes on Strix @pupposandro / @luceboxai — DFlash/PFlash/DSpark speculative engines @ciruai — gfx1151 Laguna quant recipes that hold up @dcapitella — toolboxes, benches, host tuning (the on-ramp; aka kyuz0) @hec_ovi — Compose packs with measured prefill/decode + ROCmFPX/vLLM/3D @NathanW1014 — strix-halo-vulkan llama.cpp fork @1bitlabs — multi-box research; Hermes + Laguna agent economics @PromptInjection — Strix fine-tune guides (Win/Linux) @soloish90 — ops: drivers, toolbox freeze, prefill bottleneck honesty @DevOuterReaches — Moonshine: Kimi K3 expert-streaming on one 128GB box @lhl — measurement science / Strix-vs-Spark honesty (quiet on X) @Level1Techs — Proxmox + ROCmFP4 + MTP depth guides Not really on X — GitHub instead (sorry if you are actually on X): deseven — + Homelab Discord · Hal0 — full Strix home AI appliance · julianmb — DeepSeek Flash ROCmFPX+DSpark packaging · daimonionnn — Windows ROCmFPX server · ayysasha — dual Strix vLLM over USB4 · ianbarber — Claude Code gfx1151 setup skills · urbanswelt — IOMMU-off host tuning benches · lhl repos —
Show more
spur your compute: . Multi-thousand GPU clusters now run on SPUR it can also control k0s so workloads can use k8s or spur on unified management plane. We will be ramping up on very large CPU fleets too. Give it a try, submit an issue or PR.
Show more
my team figured out how to run Kimi K3 on @AMD MI355X at 952 tok/s/node and 118 tok/s single stream. 3.8x the aggregate throughput/node and 1.3x the single stream decode of B200 and beat B300 on performance per dollar: 48 vs 33 tok/s/$ details in reply
Show more
🚨 BREAKING: these engineers figured out how to serve Kimi K3 on @AMD MI355X at 952 tok/s/node and 118 tok/s single stream! this crushes B200 by 3.8x in aggregate throughput/node and 1.3x in single stream decode + beats B300 on performance per dollar (48 vs 33 tok/s/$) See how in the thread.
Show more
We want AMD delightful to use, both for software and hardware. Great work @jackhuynh and team on the Halo box
The Ryzen Halo by @AMD is a work of art, it's quiet and cool, has a subtle but slick Need for Speed Underground aesthetic, blazing fast, most AI software worked OOB. Between this & the @FrameworkPuter desktop, AMD is shipping the most beautiful electronic devices on the market.
Show more
Great to have Day 0 support for a frontier open source model and Day 0 support of MI455 with the strong partnership with @vllm_project. Speed is the moat 🚀🚀🚀
AMD previewed MI450/455x with the community last week, and the team got K3 running on MI350/355x day 0. Speed is the moat, great work AMD team bring K3 to the world @AnushElangovan @EmadBarsoumPi Andy and Peng!
Show more
"Chance favors the prepared mind"
@__tinygrad__ @AMD @sgl_project @Kimi_Moonshot AMD, SGLang, and Moonshot shipping together out of the gate. That's not luck, it's infrastructure people preparing for this exact moment. The chip wars just got a lot more interesting.
Show more
🙏🏾🙏🏾🙏🏾Great work by the teams to make it great out of the box.
Kimi K3 running on AMD MI350X with SGLang. Amazing work to @AMD @sgl_project @Kimi_Moonshot this worked great out of the box!
🙏🏾🙏🏾🙏🏾 the credit goes to the team behind making the changes required to move fast. Speed is the moat 🚀🚀🚀
Anush and his team’s efforts made me a believer that ROCm could compete. Trust me, I was a skeptic. Anush would and still does personally respond on X publicly to every credible criticism, not with excuses or defensiveness, but with action. I really admire that and wish more tech execs would do that. I remember when ROCm, first launched in 2016, had a quarterly release schedule for major (kinda) feature releases. Now it’s every 6 weeks. Is “speed the moat”? We will see.
Show more
Thank you @MTSlive and @sophiadew for having me and talking about all the announcements at @AMD Advancing AI 2026.