Register and share your invite link to earn from video plays and referrals.

Kydo
@0xkydo
Fascinated by greatness and exploring the open frontier | Eigen Labs (Darkbloom, Yukon...)
893 Following    14.8K Followers
My prior is that Apple will be the backbone for local businesses who are increasingly worried about big tech overreach. This will be an amazing tailwind for local AI and Darkbloom's fleet utilization.
Show more
Apple connected four Mac Studios over Thunderbolt and ran a trillion-parameter AI model to find and fix a graphics coding bug. The whole setup ran from one wall outlet. Now @Apple is pitching Macs to businesses as local AI machines: buy the hardware once, then run coding agents and other AI workloads without paying per token. That’s a much more interesting challenge to the cloud than another AI feature baked into macOS.
Show more
No more recovery mode needed to turn on RDMA.
New MacOS 27 setting for clustering Macs for AI
It’s a big day for the Mac! Mac mini and Mac Studio provide a huge leap in performance, especially for AI.
0
643
27.8K
949
Forward to community
Yesterday marked one month anniversary of @DarkbloomAI moving to paid tier on OpenRouter. We started with 4B tokens a day and today we are excited to be serving 20B+ tokens a day across 8 models. We have grown our provider base to 1.1K+ active at all times and plan to expand it much higher as we expand the model base and more powerful M5 chips are getting shipped over next few weeks. You can now run Darkbloom on any Macs (including corp) without our MDM. Last week, we piloted our new higher grade privacy model which allows us to bind machines to stronger trust guarantees leading to Apple's certificates. You must be on MacOS 27 to be part of this and you will be soon required to be on OS 27+. During last 2 weeks, we have shipped a bunch of features to increase the QoL for our providers which includes Global Payouts, better routing, MLX(.)Fast inference engine improvements etc. Thank you everyone for being early to our network and supporting us through all the journey. I also want to thank the OpenRouter team (@pingToven, @shashankgoyal95, Eddie, Will, @OpenRouterMatt) for making this happen and continuing to support us through adding new models. This week we're at our company offsite + retreat planning through all the things we need to ship in the coming weeks and months. We will share a much detailed update after this. If you are on Darkbloom, make sure you're part of the Slack community.
Show more
Is Nat Friedman the 🐐 product leader? Launched AI coding with Github Copilot in <check notes> 2021/2022 Now mainstreaming AI assistants with Muse
Today is the day! Looking forward to seeing all the M5 Ultra benchmark.
The first person to put M5 Ultra on Darkbloom gets 100 bucks and bragging rights forever.
Friday is generally my favorite day because it is the end of calls and the start of two days for reading and deep wprk.
This is the first time in Darkbloom's history that we are the first provider to support a model on @OpenRouter ❤️ I couldn't have chosen a better partner for this milestone. @PrismML focuses on concentrating intelligence, so more models can run on distributed Mac machines!
Show more
Extremely excited to be hosted on Darkbloom! Amazed at the speed at which @sreeramkannan and team have made this possible
Bonsai 2 at <6GB size and ~Opus 4.6 performance. Live on Darkbloom. This breakthrough is from @PrismML founded by Caltech Prof. @BabakHassibi a leading information theorist. Optimizing intelligence per bit is an extremely valuable objective for Local AI. It is only possible for us to own our own intelligence if it will fit into the devices we own. Thats what the PrismML team has achieved here. It has only been 9 months from a frontier release (Opus 4.6) to getting it to fit in your phone! As we worry about frontier superintelligence concentrating power, we are delighted to see the counterweight emerging from open models that can fit into our phones. The darkbloom team @0xkydo and @gajesh were so excited by this breakthrough that they worked overnight to bring this to life on Darkbloom. Excited to offer the first hosted service for this model to everyone on Darkbloom, the compute grid powered by real people! Go try it on our chat or provision your Macbook to service this model for the world!
Show more
Bonsai 2 is a very capable model for daily knowledge work. Here is the essay it wrote on "the societal, political, and economic importance of AI that can run locally on any consumer computer or even a phone." You can try it out on and we are looking to support this on OpenRouter soon as well. --------- The Quiet Revolution: Why Local AI Matters - By Bonsai 2 27B For much of the past decade, artificial intelligence has been imagined as something distant and centralized: a service hosted in vast data centers, accessed through cloud platforms, and controlled by a small number of large technology companies. This model has produced remarkable capabilities, but it has also concentrated power, created new dependencies, and raised serious questions about privacy, security, and access. A different possibility is emerging: AI that can run locally on ordinary consumer devices—laptops, desktops, tablets, and even smartphones. This shift is not merely a technical improvement. It is a potentially transformative development with profound societal, political, and economic implications. Local AI matters because it changes who controls intelligence. When a model runs on a person’s own device, the user is no longer merely a customer sending data to a remote server. The user becomes, in a meaningful sense, the operator of an intelligent system. That distinction may seem small, but it reshapes the relationship between individuals, institutions, and technology. It affects how people communicate, work, learn, and participate in public life. It also affects how governments, corporations, and civil society understand power in the digital age. Societal Importance: Privacy, Autonomy, and Access The most immediate societal benefit of local AI is privacy. Many of today’s AI services require users to send prompts, documents, images, or other data to external servers. Even when companies promise strong privacy protections, the architecture itself creates risk. Data must travel across networks, be stored, processed, and potentially logged. For sensitive matters—medical questions, legal advice, personal relationships, political opinions, or confidential business information—this can be unacceptable. Local AI reduces that exposure. If a model runs on the user’s own device, the data can remain under the user’s control. A person can ask a local assistant to summarize a medical report, draft a letter, analyze a contract, or help with a personal decision without sending the underlying material to a third party. This is especially important for vulnerable populations: patients, journalists, activists, lawyers, journalists, researchers, and anyone whose privacy is essential to their safety or dignity. Local AI also supports autonomy. In a world where many digital services are designed to shape behavior, local intelligence can be a tool for personal agency rather than a channel for external influence. A user can choose which model to run, how to configure it, and what it is allowed to do. This is not a perfect solution—local models can still be biased, misleading, or poorly designed—but it creates a more transparent and controllable environment. The user can inspect the system, update it, or replace it without depending on a single provider’s terms of service. There is also a powerful accessibility dimension. Much of today’s AI is available through subscription services, cloud APIs, or high-end hardware. Local AI can lower barriers to entry. A person with a modest smartphone or an older laptop may still be able to run useful models for writing, translation, coding, research, or education. This could be especially significant in regions where internet access is unreliable, expensive, or monitored. In such contexts, local AI may function not only as a productivity tool but as a form of digital resilience. Political Importance: Sovereignty, Censorship, and Power The political importance of local AI is less obvious but equally significant. Intelligence is becoming a strategic resource. Governments, corporations, and other powerful actors are investing heavily in AI because it can influence decision-making, communication, surveillance, and economic competition. If most AI capability remains concentrated in a few cloud providers, those providers become critical nodes in the political and economic order. Their policies, outages, legal compliance, and commercial interests can shape what information is available, what speech is permitted, and what kinds of intelligence are accessible. Local AI can reduce that concentration. When intelligence can run on ordinary devices, it becomes harder for any single entity to monopolize access to advanced reasoning, translation, analysis, or creative assistance. This does not eliminate power imbalances—large companies still design many models, and governments still regulate them—but it introduces a more distributed structure. It creates space for independent actors, small businesses, public institutions, and individuals to develop their own capabilities without depending on a single platform. This has implications for censorship and control. In authoritarian contexts, centralized AI services can be easily blocked, monitored, or co-opted. Local models, by contrast, are harder to shut down remotely. A government may be able to restrict access to cloud services, but it is much more difficult to prevent people from running software on their own devices. This does not mean local AI is inherently liberating; it can also be used for surveillance or manipulation. But it does mean that the balance of power shifts. The capacity for intelligent assistance becomes less dependent on the goodwill of a central authority. Local AI also matters for democratic participation. In an era of misinformation, deepfakes, and algorithmic manipulation, citizens need tools that help them evaluate information, summarize complex policy, translate documents, and understand technical arguments. If those tools are available locally, they can support a more informed public. A voter can use a local model to compare policy proposals, a journalist can analyze leaked documents, and a community organization can translate public notices into multiple languages. The political value lies not in the model’s perfection, but in its accessibility and independence from a single commercial or state-controlled infrastructure. There is also a question of data sovereignty. In many countries, data is increasingly treated as a national asset. Local AI can help institutions keep sensitive information within their own jurisdiction. A hospital, university, or government agency can run models on its own hardware without sending patient records, research data, or classified documents to foreign servers. This is especially important in fields where confidentiality is legally or ethically required. Economic Importance: New Markets, New Work, and New Productivity Economically, local AI could be one of the most important technological shifts of the coming decades. Today, much of the AI economy is built around cloud compute, data centers, and enterprise subscriptions. That model is powerful, but it is also expensive and concentrated. Local AI changes the economics by moving intelligence closer to the user and closer to the data. For individuals, local AI can increase productivity. A student can use a local model to study, a freelancer can draft proposals, a small business owner can analyze customer feedback, and a developer can generate code. These uses may not replace human judgment, but they can reduce the cost of cognitive labor. Over time, this could lower the barrier to entry for many professional and creative activities. For businesses, local AI offers a different value proposition than cloud AI. Some tasks are better suited to local processing because they involve sensitive data, require low latency, or must work offline. A law firm can analyze contracts without sending them to a cloud service. A manufacturer can run diagnostic models on factory equipment without exposing proprietary data. A retailer can use local models to personalize customer interactions while keeping transaction data on-site. In these cases, local AI is not just a privacy feature; it is a business requirement. Local AI may also create new markets. Instead of selling only cloud subscriptions, companies can sell models, fine-tuned versions, hardware optimization, and local deployment services. Developers can build applications that run entirely on-device, creating a new category of software that is faster, more private, and less dependent on connectivity. This could be especially important for mobile applications, where users expect instant responses and do not want to upload personal data to the cloud. There are also labor-market implications. Local AI will not eliminate the need for human workers, but it will change the nature of many jobs. Tasks that once required specialized expertise may become more accessible to ordinary users. A small clinic may be able to use AI-assisted triage tools, a local government office may be able to automate routine administrative work, and a small manufacturer may be able to improve quality control without hiring a large data science team. This could increase economic productivity, but it could also create displacement pressures. The key question will be whether societies can adapt education, training, and social policy to match the new economic reality. One of the most important economic effects may be the reduction of digital inequality. Historically, advanced computing has been concentrated in wealthy nations and wealthy institutions. Local AI could help democratize access to intelligence by allowing useful models to run on devices that people already own. This does not mean that all devices will be equally capable, and there will still be a gap between high-end and low-end hardware. But the direction of travel is significant: intelligence is becoming something that can be embedded in everyday objects rather than reserved for specialized infrastructure. The Limits and Risks It would be irresponsible to present local AI as a simple solution. There are important limitations. Consumer devices have constraints on memory, power, and thermal capacity. Not every model can run efficiently on a phone, and not every task requires a large model. Smaller local models may be less capable than the largest cloud-based systems, and they may require careful design to be useful. There are also risks of misuse: local AI can be used to generate convincing misinformation, automate manipulation, or support surveillance if deployed without safeguards. There is also a governance challenge. If intelligence becomes more decentralized, how do we ensure that it is safe, reliable, and aligned with human values? How do we prevent harmful models from spreading? How do we protect users from poorly designed or deceptive systems? These questions will require new forms of regulation, technical standards, and public oversight. Local AI does not remove the need for accountability; it changes where accountability must be located. There are also environmental and energy considerations. Running AI locally can reduce the need for massive data centers, but it also increases the energy demand of consumer devices. The net environmental impact will depend on hardware efficiency, model size, usage patterns, and the energy mix of the regions where devices are used. A responsible local AI ecosystem will need to balance capability with efficiency. A New Relationship Between People and Intelligence At its core, the importance of local AI lies in a shift in relationship. For centuries, intelligence has been associated with institutions: courts, universities, governments, corporations, and specialized experts. The digital age has already begun to distribute information, but it has often done so through centralized platforms that mediate access. Local AI goes further. It places a form of intelligence directly in the hands of individuals and small organizations. This does not mean that every person will become an expert in machine learning. Most users will not need to understand the technical details. They will simply have access to a tool that can help them think, write, analyze, and create. But the fact that the tool runs locally changes the meaning of that access. It is not just a service. It is a capability. That distinction is why local AI matters so much. It has the potential to make privacy more practical, to strengthen digital sovereignty, to reduce dependence on a small number of cloud providers, and to expand economic opportunity. It can support more independent journalism, more resilient public institutions, and more empowered individuals. It can also create new forms of innovation, especially in areas where data sensitivity or connectivity limits make cloud-based AI impractical. The coming decade will likely be defined not only by how intelligent AI systems become, but by where they run. If intelligence remains concentrated in a few remote data centers, it will continue to concentrate power. If it becomes available on ordinary devices, it may help distribute power in ways that are more inclusive, more private, and more resilient. The rise of local AI is therefore not just a technical trend. It is a social and political development with the potential to reshape the relationship between people and the systems that shape their lives. -----------
Show more
Darkbloom is the first model provider to support Ternary Bonsai 2 27B -- concentrated intelligence that fits on your phone. Try it now: First 250 users get 100 million free tokens; PrismML's new flagship model: - a ternary compression of Qwen3.8 27B at 2bits per parameter. - 8.5 GB total, 5x smaller than original - keeps 98.2% of the Qwen's FP16 benchmark performance. - 75% cheaper than Qwen 27B. Qwen3.8 27B already operates comparably with Opus 4.6 and 5.6 Luna on certain tasks. This one does it in the memory of a phone. 262K context, image input, Apache 2.0. From our first run on the network, on a single M5 Max with no caching: - 35 tok/s decode at 1K context, - 31 tok/s at 10K, - 19 tok/s at 50K. But we expect more performance gain coming in a few weeks! That's a full 27B reasoning model running comfortably on any Mac. 1,000+ Macs are serving on Darkbloom right now. Go try it out!! Thank you to @BabakHassibi @SahinLale @HessianFree @rsadri_ml @tushar_bans @evaninwords and the whole PrismML team. This is exactly the kind of model Darkbloom was built for. You can read the essay by Bonsai on Why Local AI Matters:
Show more
Today's my birthday, and I could not have asked for a better gift. Bonsai 2 fits on most phones shipping today (anything >8GB). And it outperforms Opus 4.6 and 5.6 Luna -- on tasks that would have sounded absurd to attempt anywhere two years ago. I try to be disciplined about timelines. My most optimistic estimate for something like this was early-2027. It's September 2026. Huge thank you to the @prismml team -- @BabakHassibi @SahinLale @HessianFree @rsadri_ml @tushar_bans @evaninwords -- for putting this into the world. We're bringing Bonsai 2 to @DarkbloomAI tonight. The 1,000+ providers can test it right away. And whenever those machines aren't in use, they'll serve Bonsai 2 to anyone who wants to try it. So few people set up a local model themselves, and a model this good shouldn't be gated by that. More tomorrow.
Show more
Together, the MLX(.)fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon! We're ready to bring it to @DarkbloomAI: an open network of local Mac machines providing inference to the world. One thing remains: the community flagged that its license requires a separate agreement for commercial model serving, so we're holding the launch until that's in place. We believe this is a great opportunity for the local community: one where we make Qwen models faster and more accessible, and the people running them share in the value they create. .@Alibaba_Qwen @QwenDevs, we'd love to work together on this. If anyone else knows someone we can talk to, we'd love to have that conversation. Let's make Qwen 3.8 Flash on Darkbloom a reality!
Show more
The future of AI will be Owned locally Shared among communities Distributed across the world
My conversation with Michael Moritz, one of the great venture investors of the last 40 years. Michael joined Sequoia in 1986 and co-led the firm with Doug Leone (@dougleone) from 1995 to 2012. His investments include Google, Yahoo, PayPal, and Stripe. His new book, Ausländer, traces his parents' escape from Nazi Germany and helps explain his lifelong interest in what shapes exceptional people. We discuss: - The infamous Steve Jobs profile - Why Don Valentine hired him - The question he finds most revealing when interviewing people - Monomania and the cost of greatness - Why he has never felt good at anything - Writing, journalism and AI - What he learned about himself writing Ausländer I'd recommend watching this one if you can. Michael was incredibly thoughtful and reflective throughout, particularly when talking about his parents, childhood, and the more difficult parts of his personal story. Enjoy! TIMESTAMPS 0:00 Intro 0:53 Family History & Identity 7:03 Survival & Outsider Instinct 18:53 Studying Exceptional People 33:50 Self-Doubt & Success 41:17 Steve Jobs & Obsession 52:50 Joining Sequoia 1:03:22 Leadership & Alex Ferguson 1:08:25 Elon Musk & AI 1:17:12 Becoming Who You Are
Show more
0
56
998
110
Forward to community
New blog post: We had a million lines of Python to clean up. On September 2nd @Teknium asked Hermes Agent to do it. 1,393 subagents and nineteen hours later, the codebase was 34.4% smaller, saving us nearly $2m in engineering hours.
Show more
0
300
4.5K
274
Forward to community
Somehow I missed this, but we crossed 1,000 people in our Slack channel 🤯🤯 It's incredible to know this many people care. @gajesh and I truly appreciate everyone who has joined, helped others, and stayed patient with Darkbloom (still in alpha!) See you all at 2,000!
Show more
Welcome 👋 We're launching @Alibaba_Qwen 3.8-Flash-Next on BOTH and today! This is our first challenge on We see a lot of overlap between the MLX and CUDA communities, and we want to see how far both can push the same model. So we're putting their progress on the same graph for this one! Just some friendly neighborhood competition 😉 The challenge: make Qwen 3.8 Flash Next run faster on two concrete setups: --- our Qwen port of @antirez's ds4 C/CUDA engine with @UnslothAI's Qwen3.8-Flash-Next-GGUF on 1x NVIDIA DGX Spark. --- our Qwen Swift/Metal engine, built on the MLX and mlx-swift-lm forks, on 1x M5 128 GB MacBook. Both tracks measure one response stream at a time. Your changes must pass the correctness checks, then beat the reference engine on the same machine. The score combines faster prompt processing (prefill) and token generation (decode), weighted 25% and 75% in the geometric mean. We chose Qwen 3.8 Flash Next because it beats Claude Opus 4.6 Max reasoning on 94% of the benchmarks from Qwen's published comparison. That's 15 out of 16, spanning coding, reasoning, instruction following, and vision! Also this model can be ran on a single Spark or a 128 GB MacBook. That's what makes this worth pushing: every inference improvement makes that capability more useful on developer hardware. Our current starting decode rates are around 17–18 tokens/sec on the Spark and 35–36 tokens/sec on the Mac. Let's see how far we can push both! Thank you to @TheDavidTai and @GumbiiDigital for building the entire challenge (it was not easy!), @antirez @ivanfioravanti & many others for their ds4 work (the Gs), and the MLX and Qwen teams for the work we're building on. One caveat: this is a significantly larger model, so local testing will be harder for some smaller Macs. We recommend a DGX Spark or a Mac with more than 128 GB of memory to give yourself room to work.
Show more
.@googleGemma 4 26B is another success on MLX Fast! We've reached a 2.6x improvement in our combined prefill and decode benchmark. 42 solvers contributed 180 promoted improvements. Thank you to everyone who's been part of this one! What makes this one different Last week, we set out to if can be the Open R&D arm for Darkbloom. Would people still be interested if we're optimizing a model that's not on the frontier? The answer was yes. The challenge was to take eight prompts coming in at the same time and process them faster concurrently. That's the workload a Darkbloom provider needs to handle: multiple people requesting the same model at once, with each getting a usable response speed. We're now seeing 654.2 tok/s in total decode throughput, averaging 81.8 tok/s per prompt across those eight concurrent requests. For context, that's roughly the same decode speed per prompt as oMLX serving a single prompt. Getting there with eight requests running together took a lot of community work. Bringing the work into Darkbloom We're working on bringing these improvements into Darkbloom. Based on the concurrency benchmark, we expect a 4–6x improvement in concurrent throughput for Gemma as that work lands. Production results are still to come. The upstream PRs are open across our MLX, MLX Swift, model engine, and Darkbloom repos. I'll put the links below. Would love for folks to review them if they're interested. What we're improving about the challenge itself We've now run three challenges, and we've learned a lot about where the infrastructure needs work. We're taking a short pause before the next Darkbloom-focused challenge to clean that up and automate more of the engine work and the challenge process. @TheDavidTai and @Spangler3000 will help us through a lot of this. I really can't thank them enough. And thank you to everyone who's spent time finding improvements, testing them, and helping us figure out what needs to get better. What's next (1) Keep optimizing the models people already use on Darkbloom. I think this is a useful direction for take the problems we see in serving and bring the community's improvements back to providers. Before shipping the next challenge, we're probably going to take a break to clean up our infrastructure on this side. (2) Bring the same platform to other communities 😬 We don't want the machine sitting idle while we work on the infrastructure, and we have something new in mind. The MLX community has set a pretty high bar. I am very curious to see if introducing a competitive spirit into the community could create some positive chemistry. Hopefully we'll have more to share tomorrow! If you read this far, thank you so much for participating and being part of our journey. Also reply: "thank you, @TheDavidTai" for all his hard work maintaining the system and making these challenge happen!
Show more