Register and share your invite link to earn from video plays and referrals.

Sudo su
@sudoingX
GPU/local LLM. more RAM and OSS... everywhere
1.1K Following    36.3K Followers
bonsai 2 27b just built this from one paragraph of prompt in one shot, all of it out of a 5.9gb file on an rtx 3060 12gb. i did not expect frontend taste at this size. small models usually get the logic right and the layout wrong, this one got the layout right, and it thought for a long time to do it, 41k tokens over 46 minutes, the context ran out to 77k and it never lost the thread. this is a ternary compression of qwen 3.8 27b, 26 tok/s fresh on a five year old gaming gpu, 13 tok/s at 77k deep, the whole 262k window resident. for what it is, on this card, this is insane, and i cannot wait to run it through real agentic coding on hermes agent, the tool loop, the builds that break and have to recover. @PrismML keep going, this is the one that runs on the card people actually own.
Show more
i do not see enough videos of hermes agent being shared in this community and it should be the place they live. dear ai content creators, small account or big one, if you have a video of hermes agent doing real work, a build, a serve, a workflow, a phone setup, post it in this space and i will personally amplify it if the work is good. my reach is on the table for you, do not leave it there. drop your links in the comments too, i go through every one.
Show more
qwen 3.8 flash next built this landing page in 29 minutes and served it on my tailnet on its own. i am running official fp8 on 2x dgx spark, full 256k context loaded, 45 tok/s with mtp on. i have run deepseek and glm on these boxes and it was good, this one just feels right. qwen 3.8 flash next is multimodal native, a vision encoder in the same weights, so the serve that wrote this page can read the screenshot it took of it. and it stays sharp at the depth where the others start drifting. qwen 3.8 flash next it is.
Show more
hermes agent is going to explode this week btw.
i think this is the biggest hermes agent onboarding unlock local ai has gotten. the number one question i get every single day is "what model can my machine run", entire vram ladders and spreadsheets exist just to answer it, and hermes desktop made it one click. now it reads your hardware, picks the model that actually fits, pulls it, wires the runtime. this is how local ai wins, remove every step between a normal person and their first local model.
Show more
when t3 subscriber kid tries hermes agent for the first time and files a bug and nobody calls him a fool
the loudest voices in ai spent two years lobbying to ban the best chips from china. they got exactly what they asked for, and it built china a second ai industry from scratch. huawei's atlas 950 is the receipt. an exaflop of fp8 on 256tb of unified memory, all ascend silicon, because you told them they couldn't buy nvidia. eighteen months, start to finish. and the models ship open regardless, kimi k3, qwen, deepseek, all downloadable. dario was loudest who wrote the essays that shaped that ban. his company ships zero open weights, just paid $1.5 billion to settle training claude on half a million pirated books, and files to ipo in october near $965 billion. i'm not alleging a motive, i'm telling you to look at the calendar. the man who lobbied hardest to lock down compute runs the lab that profits most when open competition can't keep up.
Show more
i posted this a couple weeks ago about needing a second dgx spark, and nvidia just shipped me one. the unit and the connectx cable, on a truck to bangkok within a week of me saying it out loud here on x. sit with that for a second anon. the biggest company on earth moved that fast for one guy in a room in bangkok. no stanford lab connection behind me, no committee, no comms team drafting my posts, just a builder putting real numbers in the open where anyone can tear them apart. the polished accounts with the pedigrees have been saying "democratize ai" for years, nvidia shipped the box to the guy actually doing it. giants aren't supposed to move like this, and that's exactly why the open side feels different right now. nvidia is carrying real weight for the local ai side, backing the people who own their compute instead of renting cognition from an api, and doing it quietly, on the ground, for whoever's actually doing the work. i think that's the whole reason this movement has legs man. here's what happens next. two dgx sparks, one connectx cable, 256 gigs of unified memory, tensor parallelism across both. the models that never fit on one box, glm 5.2, nemotron, deepseek v4 flash at a million tokens, they gonna clank now. i'm going to learn a stupid amount from this setup man, and share all of it, same as always. biggest shoutout from deepest depth of my heart to @Coolmark482 for making all this happen and for seeing the local ai community as worth betting on. thank you.
Show more
a used $200 gpu runs a 27b ai agent at 128k context on your desk right now, fully offline. we are so not ready for 2027.
this is the drop the local ai crowd should be losing their minds over. poolside just dropped laguna s 2.1: 118b total parameters, only 8b active per token, a full 1m context window, open weights under a real open license, on huggingface today. look at the chart. it lands at 71 on terminal-bench at 118b, sitting above deepseek v4 pro max at a trillion params, above inkling at 1.5 trillion, above nemotron 3 ultra. it's beating models ten times its size and losing only to kimi k3, which is 24 times bigger. that's the efficiency frontier, up and to the left, exactly where you want a model to sit. but here's the part that made me sit up: it runs on a single dgx spark. and this is what nobody's saying loud enough. the dgx spark is the moe king. a dense 118b would crawl on it, the bandwidth chokes reading every weight each token. a moe with 8b active only ever reads 8b, so the spark's 128 gigs holds the whole model while generation stays fast. big brain, light footprint, the exact shape the spark was built to run. open, frontier competitive, moe efficient, and it fits on a box on your desk. that's the whole thesis in one release: you don't need a datacenter, you need the right architecture on the right hardware. go grab the link below, weights are up.
Show more
0
65
1.2K
121
Forward to community
when claude cowork bros try cursor agent window
three weeks ago america pulled fable 5 offline overnight. today reuters reports china is discussing restrictions on overseas access to its own top models. alibaba, bytedance, all in the room. closed models, open weights, even unreleased ones on the table. read that again. the only two countries producing frontier models are now both drafting rules to keep them home. and here's the part nobody wants to say out loud. china open sourced its way up. qwen, deepseek, glm, free weights flooding the world while they were behind. now glm 5.2 sits beside the frontier and suddenly beijing is talking about state assets and national security crimes. everyone loves open source until they're ahead. nations, labs, all of them. openness was never a philosophy. it was a strategy for second place. i said your ai bill should be your electricity bill. i'm adding a deadline to that. the weights on your drive are the only ai on earth that no ministry, no congress, no boardroom can reach. and if you think it stops at the models, look at where a rule like this gets enforced. the download page. huggingface is the chokepoint. maybe it never comes to that. fill your nvme like it will. one more thing. the loudest essays making the case for export controls came from the ceo of the closed lab whose own model just got export controlled. dario wrote the argument for the wall. now every wall going up leaves metered apis as the only lane left open. i don't need to guess what's in anyone's heart. the incentives are doing all the talking, and they all point at the meter. this is the worst outcome. and it was a choice. the window is closing from both sides. download while downloading is still a thing.
Show more
CHINA CONSIDERS RESTRICTING OVERSEAS ACCESS TO CUTTING-EDGE AI MODELS China’s Ministry of Commerce has led meetings over the past month with major AI companies, including Alibaba, ByteDance, and to discuss measures that would restrict overseas access to cutting-edge AI models, including models that have not yet been released. The discussions reportedly include not only closed-source models but also open-weight models. However, the scope of application is still under debate, and the rules may ultimately apply only to future frontier models. Officials have also discussed designating the leakage or theft of proprietary AI technologies as a national security crime, with stronger penalties, as well as restricting the types of foreign capital that can invest in Chinese AI startups. The backdrop is the U.S. move to strengthen export controls on AI models, along with national security concerns over cutting-edge models that could possess advanced cyberattack capabilities. Chinese authorities are reportedly concerned that advanced U.S. cybersecurity AI models could be used to exploit vulnerabilities in Chinese software. Since the beginning of this year, China has continued to tighten measures to prevent AI technology from being transferred overseas. Authorities have investigated whether Chinese AI startups that relocated abroad violated export control laws, while also strengthening oversight of overseas transactions involving Chinese investors, technology, data, and national security concerns. Future regulations could take the form of a tiered framework based on technological capability. Basic open-source AI models may be managed through a filing system, high-performance models may be subject to security reviews, and the most sensitive frontier models may be banned from public release or restricted to use within China.
Show more
see this dip? that's july 5th. the day i flew to another city and got engaged. impressions dropped 80% because i chose the most important day of my life over the timeline. here's what nobody tells you about x monetization: the algorithm has no memory of your reasons. it only knows your reps. so now i climb back. in public. in real time. wedding is in February 24. i'm funding it with this account and my work. every 2 weeks x pays out, and i'm going to show you the climb, payout by payout, post by post. not tips from a guru. receipts from a man with a deadline. the siege starts now. watch the chart.
Show more
i got engaged this week ❤️. want to tell you what i learned. i lost my father at 11. became the provider at 13. for 15 years it was just me, the work, and the people counting on me. no degree. no safety net. the internet became my university and shipping became my religion. this week i flew to her city with 4 of my elders, the men who stood where my father would have stood. we did it the old way. flowers. sweets. blessings. and when her family asked what i do, i just smiled and said "i build software". they don't know about the DMs from companies you'd recognize. they don't need to. let your work be discovered, never announced. she said yes. this photo is me at the airport the next day, everyone scrolling, me running agents, pushing code from the gate. because the man she said yes to is a builder, and builders don't stop being builders on the good days. the good days are made OF the building. wedding in january. launching my platform in spring. everything i've built was the trailer. locked in for the greatest 6 months of my life. watch anon.
Show more
i got engaged this week ❤️. want to tell you what i learned. i lost my father at 11. became the provider at 13. for 15 years it was just me, the work, and the people counting on me. no degree. no safety net. the internet became my university and shipping became my religion. this week i flew to her city with 4 of my elders, the men who stood where my father would have stood. we did it the old way. flowers. sweets. blessings. and when her family asked what i do, i just smiled and said "i build software". they don't know about the DMs from companies you'd recognize. they don't need to. let your work be discovered, never announced. she said yes. this photo is me at the airport the next day, everyone scrolling, me running agents, pushing code from the gate. because the man she said yes to is a builder, and builders don't stop being builders on the good days. the good days are made OF the building. wedding in january. launching my platform in spring. everything i've built was the trailer. locked in for the greatest 6 months of my life. watch anon.
Show more
anon. if you want into local ai and don't know where to start, here it is. grab a used rtx 3090. six years old, 24gb of vram, still the best value per dollar in the game. load qwen3.6 27b dense at q4. that pairing is the king of the card, and i'll stand on that. that one setup gets you a real build companion and retires a subscription or two. your ai bill becomes your electricity bill.
Show more
what was the ram in your first computer, and what's in your current one?
many won’t understand, at least not yet, but local ai is required for humanity survival.
i am back. so so so fucking back man.
what was the ram in your first computer, and what's in your current one?