Register and share your invite link to earn from video plays and referrals.

Darkbloom
@darkbloomai
Open Inference Network of Idle Macs. Public Alpha. Live on OpenRouter. An @EigenLabs Project.
5 Following    10.1K Followers
We’re live on OpenRouter, Thank you @pingToven for the quick response and activation of the model!!
This is the first time in Darkbloom's history that we are the first provider to support a model on @OpenRouter ❤️ I couldn't have chosen a better partner for this milestone. @PrismML focuses on concentrating intelligence, so more models can run on distributed Mac machines!
Show more
Extremely excited to be hosted on Darkbloom! Amazed at the speed at which @sreeramkannan and team have made this possible
NEW: Bonsai 2 is LIVE on Darkbloom PrismML compressed Qwen3.8 27B to 8.5 GB and kept 98.2% of its benchmark scores. A full 27B reasoning model that fits on a phone. On Darkbloom, it runs from a browser tab. First 250 users get 100 million free tokens.
Show more
Fantastic! Try Bonsai 2 27B on Darkbloom 🌳🌳
Bonsai 2 is a very capable model for daily knowledge work. Here is the essay it wrote on "the societal, political, and economic importance of AI that can run locally on any consumer computer or even a phone." You can try it out on and we are looking to support this on OpenRouter soon as well. --------- The Quiet Revolution: Why Local AI Matters - By Bonsai 2 27B For much of the past decade, artificial intelligence has been imagined as something distant and centralized: a service hosted in vast data centers, accessed through cloud platforms, and controlled by a small number of large technology companies. This model has produced remarkable capabilities, but it has also concentrated power, created new dependencies, and raised serious questions about privacy, security, and access. A different possibility is emerging: AI that can run locally on ordinary consumer devices—laptops, desktops, tablets, and even smartphones. This shift is not merely a technical improvement. It is a potentially transformative development with profound societal, political, and economic implications. Local AI matters because it changes who controls intelligence. When a model runs on a person’s own device, the user is no longer merely a customer sending data to a remote server. The user becomes, in a meaningful sense, the operator of an intelligent system. That distinction may seem small, but it reshapes the relationship between individuals, institutions, and technology. It affects how people communicate, work, learn, and participate in public life. It also affects how governments, corporations, and civil society understand power in the digital age. Societal Importance: Privacy, Autonomy, and Access The most immediate societal benefit of local AI is privacy. Many of today’s AI services require users to send prompts, documents, images, or other data to external servers. Even when companies promise strong privacy protections, the architecture itself creates risk. Data must travel across networks, be stored, processed, and potentially logged. For sensitive matters—medical questions, legal advice, personal relationships, political opinions, or confidential business information—this can be unacceptable. Local AI reduces that exposure. If a model runs on the user’s own device, the data can remain under the user’s control. A person can ask a local assistant to summarize a medical report, draft a letter, analyze a contract, or help with a personal decision without sending the underlying material to a third party. This is especially important for vulnerable populations: patients, journalists, activists, lawyers, journalists, researchers, and anyone whose privacy is essential to their safety or dignity. Local AI also supports autonomy. In a world where many digital services are designed to shape behavior, local intelligence can be a tool for personal agency rather than a channel for external influence. A user can choose which model to run, how to configure it, and what it is allowed to do. This is not a perfect solution—local models can still be biased, misleading, or poorly designed—but it creates a more transparent and controllable environment. The user can inspect the system, update it, or replace it without depending on a single provider’s terms of service. There is also a powerful accessibility dimension. Much of today’s AI is available through subscription services, cloud APIs, or high-end hardware. Local AI can lower barriers to entry. A person with a modest smartphone or an older laptop may still be able to run useful models for writing, translation, coding, research, or education. This could be especially significant in regions where internet access is unreliable, expensive, or monitored. In such contexts, local AI may function not only as a productivity tool but as a form of digital resilience. Political Importance: Sovereignty, Censorship, and Power The political importance of local AI is less obvious but equally significant. Intelligence is becoming a strategic resource. Governments, corporations, and other powerful actors are investing heavily in AI because it can influence decision-making, communication, surveillance, and economic competition. If most AI capability remains concentrated in a few cloud providers, those providers become critical nodes in the political and economic order. Their policies, outages, legal compliance, and commercial interests can shape what information is available, what speech is permitted, and what kinds of intelligence are accessible. Local AI can reduce that concentration. When intelligence can run on ordinary devices, it becomes harder for any single entity to monopolize access to advanced reasoning, translation, analysis, or creative assistance. This does not eliminate power imbalances—large companies still design many models, and governments still regulate them—but it introduces a more distributed structure. It creates space for independent actors, small businesses, public institutions, and individuals to develop their own capabilities without depending on a single platform. This has implications for censorship and control. In authoritarian contexts, centralized AI services can be easily blocked, monitored, or co-opted. Local models, by contrast, are harder to shut down remotely. A government may be able to restrict access to cloud services, but it is much more difficult to prevent people from running software on their own devices. This does not mean local AI is inherently liberating; it can also be used for surveillance or manipulation. But it does mean that the balance of power shifts. The capacity for intelligent assistance becomes less dependent on the goodwill of a central authority. Local AI also matters for democratic participation. In an era of misinformation, deepfakes, and algorithmic manipulation, citizens need tools that help them evaluate information, summarize complex policy, translate documents, and understand technical arguments. If those tools are available locally, they can support a more informed public. A voter can use a local model to compare policy proposals, a journalist can analyze leaked documents, and a community organization can translate public notices into multiple languages. The political value lies not in the model’s perfection, but in its accessibility and independence from a single commercial or state-controlled infrastructure. There is also a question of data sovereignty. In many countries, data is increasingly treated as a national asset. Local AI can help institutions keep sensitive information within their own jurisdiction. A hospital, university, or government agency can run models on its own hardware without sending patient records, research data, or classified documents to foreign servers. This is especially important in fields where confidentiality is legally or ethically required. Economic Importance: New Markets, New Work, and New Productivity Economically, local AI could be one of the most important technological shifts of the coming decades. Today, much of the AI economy is built around cloud compute, data centers, and enterprise subscriptions. That model is powerful, but it is also expensive and concentrated. Local AI changes the economics by moving intelligence closer to the user and closer to the data. For individuals, local AI can increase productivity. A student can use a local model to study, a freelancer can draft proposals, a small business owner can analyze customer feedback, and a developer can generate code. These uses may not replace human judgment, but they can reduce the cost of cognitive labor. Over time, this could lower the barrier to entry for many professional and creative activities. For businesses, local AI offers a different value proposition than cloud AI. Some tasks are better suited to local processing because they involve sensitive data, require low latency, or must work offline. A law firm can analyze contracts without sending them to a cloud service. A manufacturer can run diagnostic models on factory equipment without exposing proprietary data. A retailer can use local models to personalize customer interactions while keeping transaction data on-site. In these cases, local AI is not just a privacy feature; it is a business requirement. Local AI may also create new markets. Instead of selling only cloud subscriptions, companies can sell models, fine-tuned versions, hardware optimization, and local deployment services. Developers can build applications that run entirely on-device, creating a new category of software that is faster, more private, and less dependent on connectivity. This could be especially important for mobile applications, where users expect instant responses and do not want to upload personal data to the cloud. There are also labor-market implications. Local AI will not eliminate the need for human workers, but it will change the nature of many jobs. Tasks that once required specialized expertise may become more accessible to ordinary users. A small clinic may be able to use AI-assisted triage tools, a local government office may be able to automate routine administrative work, and a small manufacturer may be able to improve quality control without hiring a large data science team. This could increase economic productivity, but it could also create displacement pressures. The key question will be whether societies can adapt education, training, and social policy to match the new economic reality. One of the most important economic effects may be the reduction of digital inequality. Historically, advanced computing has been concentrated in wealthy nations and wealthy institutions. Local AI could help democratize access to intelligence by allowing useful models to run on devices that people already own. This does not mean that all devices will be equally capable, and there will still be a gap between high-end and low-end hardware. But the direction of travel is significant: intelligence is becoming something that can be embedded in everyday objects rather than reserved for specialized infrastructure. The Limits and Risks It would be irresponsible to present local AI as a simple solution. There are important limitations. Consumer devices have constraints on memory, power, and thermal capacity. Not every model can run efficiently on a phone, and not every task requires a large model. Smaller local models may be less capable than the largest cloud-based systems, and they may require careful design to be useful. There are also risks of misuse: local AI can be used to generate convincing misinformation, automate manipulation, or support surveillance if deployed without safeguards. There is also a governance challenge. If intelligence becomes more decentralized, how do we ensure that it is safe, reliable, and aligned with human values? How do we prevent harmful models from spreading? How do we protect users from poorly designed or deceptive systems? These questions will require new forms of regulation, technical standards, and public oversight. Local AI does not remove the need for accountability; it changes where accountability must be located. There are also environmental and energy considerations. Running AI locally can reduce the need for massive data centers, but it also increases the energy demand of consumer devices. The net environmental impact will depend on hardware efficiency, model size, usage patterns, and the energy mix of the regions where devices are used. A responsible local AI ecosystem will need to balance capability with efficiency. A New Relationship Between People and Intelligence At its core, the importance of local AI lies in a shift in relationship. For centuries, intelligence has been associated with institutions: courts, universities, governments, corporations, and specialized experts. The digital age has already begun to distribute information, but it has often done so through centralized platforms that mediate access. Local AI goes further. It places a form of intelligence directly in the hands of individuals and small organizations. This does not mean that every person will become an expert in machine learning. Most users will not need to understand the technical details. They will simply have access to a tool that can help them think, write, analyze, and create. But the fact that the tool runs locally changes the meaning of that access. It is not just a service. It is a capability. That distinction is why local AI matters so much. It has the potential to make privacy more practical, to strengthen digital sovereignty, to reduce dependence on a small number of cloud providers, and to expand economic opportunity. It can support more independent journalism, more resilient public institutions, and more empowered individuals. It can also create new forms of innovation, especially in areas where data sensitivity or connectivity limits make cloud-based AI impractical. The coming decade will likely be defined not only by how intelligent AI systems become, but by where they run. If intelligence remains concentrated in a few remote data centers, it will continue to concentrate power. If it becomes available on ordinary devices, it may help distribute power in ways that are more inclusive, more private, and more resilient. The rise of local AI is therefore not just a technical trend. It is a social and political development with the potential to reshape the relationship between people and the systems that shape their lives. -----------
Show more
Darkbloom is the first model provider to support Ternary Bonsai 2 27B -- concentrated intelligence that fits on your phone. Try it now: First 250 users get 100 million free tokens; PrismML's new flagship model: - a ternary compression of Qwen3.8 27B at 2bits per parameter. - 8.5 GB total, 5x smaller than original - keeps 98.2% of the Qwen's FP16 benchmark performance. - 75% cheaper than Qwen 27B. Qwen3.8 27B already operates comparably with Opus 4.6 and 5.6 Luna on certain tasks. This one does it in the memory of a phone. 262K context, image input, Apache 2.0. From our first run on the network, on a single M5 Max with no caching: - 35 tok/s decode at 1K context, - 31 tok/s at 10K, - 19 tok/s at 50K. But we expect more performance gain coming in a few weeks! That's a full 27B reasoning model running comfortably on any Mac. 1,000+ Macs are serving on Darkbloom right now. Go try it out!! Thank you to @BabakHassibi @SahinLale @HessianFree @rsadri_ml @tushar_bans @evaninwords and the whole PrismML team. This is exactly the kind of model Darkbloom was built for. You can read the essay by Bonsai on Why Local AI Matters:
Show more
Last night @PrismML shipped Bonsai 2 and @darkbloomai is the first model provider to support it. We're offering the first 250 users 100m free tokens test it out. At 9x smaller than Qwen 3.8 with 98.2% of the benchmark performance. It outperforms Opus 4.6 and 5.6 Luna and its the first model that can run on the lightest weight personal MacBooks in our Darkbloom fleet. The race for bigger centralized models misses half the point. Models like Bonsai are how AI becomes ubiquitous and Darkbloom ensures the spoils of the inference economy are accessible to everyone.
Show more
Bonsai 2 at <6GB size and ~Opus 4.6 performance. Live on Darkbloom. This breakthrough is from @PrismML founded by Caltech Prof. @BabakHassibi a leading information theorist. Optimizing intelligence per bit is an extremely valuable objective for Local AI. It is only possible for us to own our own intelligence if it will fit into the devices we own. Thats what the PrismML team has achieved here. It has only been 9 months from a frontier release (Opus 4.6) to getting it to fit in your phone! As we worry about frontier superintelligence concentrating power, we are delighted to see the counterweight emerging from open models that can fit into our phones. The darkbloom team @0xkydo and @gajesh were so excited by this breakthrough that they worked overnight to bring this to life on Darkbloom. Excited to offer the first hosted service for this model to everyone on Darkbloom, the compute grid powered by real people! Go try it on our chat or provision your Macbook to service this model for the world!
Show more
Darkbloom is the first model provider to support Ternary Bonsai 2 27B -- concentrated intelligence that fits on your phone. Try it now: First 250 users get 100 million free tokens; PrismML's new flagship model: - a ternary compression of Qwen3.8 27B at 2bits per parameter. - 8.5 GB total, 5x smaller than original - keeps 98.2% of the Qwen's FP16 benchmark performance. - 75% cheaper than Qwen 27B. Qwen3.8 27B already operates comparably with Opus 4.6 and 5.6 Luna on certain tasks. This one does it in the memory of a phone. 262K context, image input, Apache 2.0. From our first run on the network, on a single M5 Max with no caching: - 35 tok/s decode at 1K context, - 31 tok/s at 10K, - 19 tok/s at 50K. But we expect more performance gain coming in a few weeks! That's a full 27B reasoning model running comfortably on any Mac. 1,000+ Macs are serving on Darkbloom right now. Go try it out!! Thank you to @BabakHassibi @SahinLale @HessianFree @rsadri_ml @tushar_bans @evaninwords and the whole PrismML team. This is exactly the kind of model Darkbloom was built for. You can read the essay by Bonsai on Why Local AI Matters:
Show more
> be @darkbloomai >gather everyday people's Macs >connect them to a grid >create a private "local AI as a service" >optimize models like @googlegemma @Alibaba_Qwen and @nvidia nemotron with your sister team @yukonresearch >run them 2-3x faster and 50% cheaper by stripping out datacenter margins >partner with @OpenRouter for distribution >pass earnings along to passionate, growing local AI community that offer up their spare compute. rinse. lather. repeat
Show more
Bonsai 2 coming to Darkbloom A 27B model in 5.9 GB. Most Macs already on the network can serve it, no new hardware needed. Stay tuned.
Today's my birthday, and I could not have asked for a better gift. Bonsai 2 fits on most phones shipping today (anything >8GB). And it outperforms Opus 4.6 and 5.6 Luna -- on tasks that would have sounded absurd to attempt anywhere two years ago. I try to be disciplined about timelines. My most optimistic estimate for something like this was early-2027. It's September 2026. Huge thank you to the @prismml team -- @BabakHassibi @SahinLale @HessianFree @rsadri_ml @tushar_bans @evaninwords -- for putting this into the world. We're bringing Bonsai 2 to @DarkbloomAI tonight. The 1,000+ providers can test it right away. And whenever those machines aren't in use, they'll serve Bonsai 2 to anyone who wants to try it. So few people set up a local model themselves, and a model this good shouldn't be gated by that. More tomorrow.
Show more
Today's my birthday, and I could not have asked for a better gift. Bonsai 2 fits on most phones shipping today (anything >8GB). And it outperforms Opus 4.6 and 5.6 Luna -- on tasks that would have sounded absurd to attempt anywhere two years ago. I try to be disciplined about timelines. My most optimistic estimate for something like this was early-2027. It's September 2026. Huge thank you to the @prismml team -- @BabakHassibi @SahinLale @HessianFree @rsadri_ml @tushar_bans @evaninwords -- for putting this into the world. We're bringing Bonsai 2 to @DarkbloomAI tonight. The 1,000+ providers can test it right away. And whenever those machines aren't in use, they'll serve Bonsai 2 to anyone who wants to try it. So few people set up a local model themselves, and a model this good shouldn't be gated by that. More tomorrow.
Show more
Together, the MLX(.)fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon! We're ready to bring it to @DarkbloomAI: an open network of local Mac machines providing inference to the world. One thing remains: the community flagged that its license requires a separate agreement for commercial model serving, so we're holding the launch until that's in place. We believe this is a great opportunity for the local community: one where we make Qwen models faster and more accessible, and the people running them share in the value they create. .@Alibaba_Qwen @QwenDevs, we'd love to work together on this. If anyone else knows someone we can talk to, we'd love to have that conversation. Let's make Qwen 3.8 Flash on Darkbloom a reality!
Show more
Somehow I missed this, but we crossed 1,000 people in our Slack channel 🤯🤯 It's incredible to know this many people care. @gajesh and I truly appreciate everyone who has joined, helped others, and stayed patient with Darkbloom (still in alpha!) See you all at 2,000!
Show more
.@googleGemma 4 26B is another success on MLX Fast! We've reached a 2.6x improvement in our combined prefill and decode benchmark. 42 solvers contributed 180 promoted improvements. Thank you to everyone who's been part of this one! What makes this one different Last week, we set out to if can be the Open R&D arm for Darkbloom. Would people still be interested if we're optimizing a model that's not on the frontier? The answer was yes. The challenge was to take eight prompts coming in at the same time and process them faster concurrently. That's the workload a Darkbloom provider needs to handle: multiple people requesting the same model at once, with each getting a usable response speed. We're now seeing 654.2 tok/s in total decode throughput, averaging 81.8 tok/s per prompt across those eight concurrent requests. For context, that's roughly the same decode speed per prompt as oMLX serving a single prompt. Getting there with eight requests running together took a lot of community work. Bringing the work into Darkbloom We're working on bringing these improvements into Darkbloom. Based on the concurrency benchmark, we expect a 4–6x improvement in concurrent throughput for Gemma as that work lands. Production results are still to come. The upstream PRs are open across our MLX, MLX Swift, model engine, and Darkbloom repos. I'll put the links below. Would love for folks to review them if they're interested. What we're improving about the challenge itself We've now run three challenges, and we've learned a lot about where the infrastructure needs work. We're taking a short pause before the next Darkbloom-focused challenge to clean that up and automate more of the engine work and the challenge process. @TheDavidTai and @Spangler3000 will help us through a lot of this. I really can't thank them enough. And thank you to everyone who's spent time finding improvements, testing them, and helping us figure out what needs to get better. What's next (1) Keep optimizing the models people already use on Darkbloom. I think this is a useful direction for take the problems we see in serving and bring the community's improvements back to providers. Before shipping the next challenge, we're probably going to take a break to clean up our infrastructure on this side. (2) Bring the same platform to other communities 😬 We don't want the machine sitting idle while we work on the infrastructure, and we have something new in mind. The MLX community has set a pretty high bar. I am very curious to see if introducing a competitive spirit into the community could create some positive chemistry. Hopefully we'll have more to share tomorrow! If you read this far, thank you so much for participating and being part of our journey. Also reply: "thank you, @TheDavidTai" for all his hard work maintaining the system and making these challenge happen!
Show more
Darkbloom just hit 500 stars on GitHub.
Darkbloom: Day 15 We have been locked in for last one week; It has been very exciting and we shipped a lot of improvements to the system. Metrics: - 16.3B tokens served: New ATH & 60% higher than last day. - $420K Network ARR (4x in 2 weeks); - 99.96% uptime (very happy considering we're a distributed network). - ~15-17% utilization at peak; median: 13% utilization. Product: - One big win: we take up 30% of traffic for Gemma 4 26B - We cut latency by ~1.5 second. We plan to ship few more features that would reduce latency by 30-90%. This will include cache-aware routing. - We shipped new routing and system optimizations which were key reason for us to cut down 429s. More such to come. - We will be adding 3-4 more models in the coming week. Excited for @Spangler3000 to be helping us with this! - We conducted MLX(.)Fast which has improved Gemma 4's performance by 2.5x -- the improvements are in process of being imported by @TheDavidTai. - We are exploring payouts solutions such as Stripe Global Payouts to add 67 more countries. - Astra has redesigned our Console UI (this is AGI, ngl). We will also be starting up our efforts for the MacOS app for better user experience. - We have also been looking more deep into Autopilot where we load up models according to the traffic and other factors to maximise the provider earnings. - Our new landing page is almost ready and coming up with one suprise ;) -- huge kudus to @0xkydo and Estella. - Also huge thanks to @iLLRiPHiTTeR for helping us while onboarding new models on OpenRouter. We were able to find out minor issues in our inference engine. - We are also working through engineering practices and cleaning up code. Thank you everyone for being part of this journey and reading along this update. We build these systems to bring people together to contribute to this AI economy. As always, it has been rewarding to keep doing what we do. Until next time.
Show more
I'll share a more detailed update on Darkbloom tomorrow. We have been locked in and shipped a lot over last few days. I'm so pumped.
Hey folks, looking for a Mac hosting provider, or an intro to one! We're looking to rent 10 x M5 Max (128GB) Macs for 3 months. Here are our specific requirements: - Bare metal with native MLX GPU access
 - Sudo access
 - Great on-call support - Full fan control + CPU/GPU temperature readings We're building where developers come to improve inference performance on Apple Silicon (we've had +100-250% performance increase for every model we worked on so far). We're also happy to feature our hosting partner's and link on the site, refer developers their way. If you know someone who can help? Please tag them below my DMs are open too!
Show more