登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Kydo
@0xkydo
Fascinated by greatness and exploring the open frontier | Eigen Labs (Darkbloom, Yukon...)
参加 August 2021
893 フォロー中    14.8K ファン
.@googleGemma 4 26B is another success on MLX Fast! We've reached a 2.6x improvement in our combined prefill and decode benchmark. 42 solvers contributed 180 promoted improvements. Thank you to everyone who's been part of this one! What makes this one different Last week, we set out to if can be the Open R&D arm for Darkbloom. Would people still be interested if we're optimizing a model that's not on the frontier? The answer was yes. The challenge was to take eight prompts coming in at the same time and process them faster concurrently. That's the workload a Darkbloom provider needs to handle: multiple people requesting the same model at once, with each getting a usable response speed. We're now seeing 654.2 tok/s in total decode throughput, averaging 81.8 tok/s per prompt across those eight concurrent requests. For context, that's roughly the same decode speed per prompt as oMLX serving a single prompt. Getting there with eight requests running together took a lot of community work. Bringing the work into Darkbloom We're working on bringing these improvements into Darkbloom. Based on the concurrency benchmark, we expect a 4–6x improvement in concurrent throughput for Gemma as that work lands. Production results are still to come. The upstream PRs are open across our MLX, MLX Swift, model engine, and Darkbloom repos. I'll put the links below. Would love for folks to review them if they're interested. What we're improving about the challenge itself We've now run three challenges, and we've learned a lot about where the infrastructure needs work. We're taking a short pause before the next Darkbloom-focused challenge to clean that up and automate more of the engine work and the challenge process. @TheDavidTai and @Spangler3000 will help us through a lot of this. I really can't thank them enough. And thank you to everyone who's spent time finding improvements, testing them, and helping us figure out what needs to get better. What's next (1) Keep optimizing the models people already use on Darkbloom. I think this is a useful direction for take the problems we see in serving and bring the community's improvements back to providers. Before shipping the next challenge, we're probably going to take a break to clean up our infrastructure on this side. (2) Bring the same platform to other communities 😬 We don't want the machine sitting idle while we work on the infrastructure, and we have something new in mind. The MLX community has set a pretty high bar. I am very curious to see if introducing a competitive spirit into the community could create some positive chemistry. Hopefully we'll have more to share tomorrow! If you read this far, thank you so much for participating and being part of our journey. Also reply: "thank you, @TheDavidTai" for all his hard work maintaining the system and making these challenge happen!
もっと見る