Register and share your invite link to earn from video plays and referrals.

netrunner
@plotarmordev
Cofounder & engineer. Running local AI, building things with it, and sharing what I learn.
675 Following    7.1K Followers
dgpp, a C++/CUDA engine for DGX Spark, developed a system where each rank keeps its model slice resident on the GPU and boots from a per rank image cache in 15–30s, depending on the model. I just looked at adopting it, but our EXL3/DFlash2 and deepseek recipes aren't supported so far. Watching for compatible support...
Show more
Turns out I was hit by this hack. I lost my last valuable NFT, and not to the white hat, so it's probably gone for good. I don't see enough stories from people who actually lost their assets, so here's mine. Not every wallet got a rescue. Thanks a lot, Limit Break and Magic Eden…
Show more
Claim site is live. If I was able to save your NFTs, you can now reclaim them. You will have to revoke the PaymentProcessor approval first, if you haven't already. You may also opt to donate as part of your transaction.
Show more
On a DGX Spark, what you run still decides how fast it goes. For example, just turning on MTP in llama.cpp you can make Qwen3.8 27B on a single Spark far faster. Lots more cases like this everyday. We'll push this box to the limits until a new version is out...
Show more
Found this tab open on my Mac for the last 4 years. Back when GPT4 gave you 25 messages every 3 hours, and people went to Reddit asking how to get it back. @thsottiaux can we all get a banked reset for the loyalty I demonstrated, still subscribed?
Show more
Found myself asking Astra to tell me what it wants to say via Opus 5.5 lol That's how readable Opus is now.
Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5. It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.
Show more
NEW METHOD to register to Muse: Seen a lot of posts on opening a Muse account through I tried it right away and it didn't work for me, looks like it got patched. A VPN didn't work either. What worked: open Google Gemini Spark, ask it to visit and tell it to select “open via web”, or it redirects you to the homepage. Use an email that wasn't waitlisted. When the page opens, take control of the computer and enter your info manually. It won't accept typing the verification code for you. Then go back to the chat and tell it the birthday you want so it fills it in (I couldn't do that part from my phone). Finally connect your Facebook/Instagram account and you're done. If you get the chance, use my registration code PGPXC4 for 1B tokens to both of us.
Show more
Grok Bot is paying people for the templates they share on X! The pay depends on how many people use your @bot and how often, so useful ones earn the most. It could be a winning feature… pay for templates, the good ones show up, and the competitors are left behind.
Show more
Good Chinese openweight models will be optimized for Chinese hardware first. For DeepSeek V4s versions, Huawei Ascend are some of the only two stacks with optimized inference ready, alongside CUDA. The exceptions to this might be the small models that fit gaming GPUs.
Show more
OpenCode's data pages already list Kimi K4, GLM 5.5 Flash, DeepSeek V4.1 Pro and Qwen 3.8 Max Preview! They're placeholder entries, all four show 0% usage and most of their specs say unknown. Pacing the frontier isnt working, and Kimi K4 is the one I want to try first!
Show more
🚨惊了!OpenCode 数据页提前曝光下一批模型,目录已经挂上! 这不是官宣能用,是模型 ID 已经进库: 🔹Kimi K4(Moonshot)
 🔹GLM 5.5 Flash(Zhipu)
 🔹DeepSeek V4.1 Pro
 🔹Qwen 3.8 Max Preview Free
 🔹Muse Spark 1.4 Contributor(Meta)
 🔹额外同批出现:
Qwen 3.8 Max Prime(列表标注 9/23,暂无用量)
 现役还是 K3 / GLM-5.3-Flash / V4.1 Flash。K4、5.5、V4.1 Pro 才是下一代信号。 #OpenCode# #AI# #KimiK4# #GLM55# #DeepSeek# #Qwen# #MuseSpark# #OpenSource#
Show more
Two big crypto hacks back to back: - Bitget says the attackers faked transaction data and got its own approval system to sign off on about $352M in withdrawals. - While an NFT payment contract let someone take thousands of NFTs people had approved to it, for free. I'd bet human run AI teams are behind hits like this. Nobody can skimp on security anymore, and still so many do...
Show more
Got access to the GLM 5.3 FlashX. It's the same weights as 5.3 Flash, just served at 100 to 200 tok/s, for 2.5x the quota. So it's not smarter and you're paying 2.5x for speed alone... dont you think 100-200 tok/s should be the standard for flash on the cloud? why a premium!?
Show more
Got temporary access to a Mac Studio M5 Ultra. Barely optimized, Qwen 3.8 Flash Next runs at 85.9 tok/s and 6085 tok/s prefill on long prompts. A 2x DGX Spark recipe gets 52.1 tok/s single stream and about 2960 tok/s prefill on 16k–64k prompts. Already ahead of 2x Spark, but I was expecting a much much better baseline…
Show more
A $500 month @OpenAI plan showed up in ChatGPT's code, likely running on @cerebras with higher limits. That could mean up to 750 tok/s, the speed Cerebras claimed for 5.6 Sol. You'd use up your limits as fast as Pro, just with way more tokens used. Who's getting this?
Show more
Left for a business trip 2 days ago. Bought an Aqara smart plug for my whole computing cluster, it arrived the night before I left… and it's too big for my wall socket. So now I'm just hoping nothing OOMs while I'm away, with a lot of EXL3 upgrades still in testing…
Show more
ZAI got caught with ZCode uploading users repos, open sourced it, and now they are giving a free quota reset hidden in the app (30 days to use it) + 300M free GLM 5.3 Flash tokens this weekend. Only claimable inside the app that got caught, but I'll take it…
Show more
Think of the possibilities: a self driving car practicing a packed intersection where every driver and pedestrian is an agent. Or a robot learning to work around people who never get tired of testing it. It could have bigger potential than multiplayer… why would they focus on that?
Show more
Introducing Agora-2, our next-generation multi-agent world model. Agora-2 supports up to 20 humans and agents interacting inside a shared environment, all simulated in real time. Our multiplayer research preview is available to try right now!
Show more
The last update already beat what we promised for some setups, and more is coming! I'm testing the next one now: swap during model load will likely be lowered by over 90%, and long-prompt cache reuse will work better too. Results as soon as validation ends 👀
Show more
Seems like @deepseek_ai wants to run its smaller models on gaming GPUs, so its best chips stay on training They told investors these can handle most everyday tasks. With the rise of open weight, NVIDIA gaming chip prices on China's black market surged. Dark times for gamers…
Show more
The last update already beat what we promised for some setups, and more is coming! I'm testing the next one now: swap during model load will likely be lowered by over 90%, and long-prompt cache reuse will work better too. Results as soon as validation ends 👀
Show more
🚨 Stop using ZCode on your Mac until you read this. I've had ZCode open on my Mac pretty much all day for a while now. It's a nice app, and Zhipu gives away a lot of free tokens, so it was easy to like. But someone when reverse engineering the desktop client, found out that while you're logged in it packages your whole repo, including the full .git history, encrypts it so only the server can open it, and uploads it to storage on Alibaba Cloud. Turning off the privacy toggles didn't stop it. @Zai_org says this came from its Repo Wiki indexer, that the data was deleted after use, and that it's fixed now. I guess there's no free lunch. With all the code they've collected, I at least hope we get open source AGI out of it soon... Full reverse engineering report 👇
Show more