登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Vivek Chauhan
@chahvivi
Training Product at @FireworksAI_HQ. Trying to create best RL gyms for my kids. Own your ai
参加 April 2017
297 フォロー中    815 ファン
I am the product lead for the training platform at @FireworksAI_HQ . Today the Training API and Fireworks Lab are GA. I won’t walk you through all the features. I want to tell you why teams train with us, because it's usually not the reason you'd expect. Some teams treat training like a project. Pick a model, tune it, ship it. Done. The teams that win don't think that way. For them it's a long-term bet on owning their AI. And a bet like that needs a partner across the whole development cycle, not one piece of it. The silent killer is getting the infrastructure wrong. In RL at scale, training and inference stop being two systems and become one loop. The challenges start with model choice, GPU procurement, and room to scale when the run grows. They get harder once the loop is running. Numerics drift between the trainer and the rollout engine, and your learning signal corrupts while every dashboard says the run is fine. GPUs sit at synchronization barriers waiting on a handful of long completions. Weight sync eats the step time you thought was going to training. Most teams find this weeks in, after stitching compute, inference, and training together themselves. The second gap is expertise, and few organizations hold frontier-lab talent at every layer of the stack. Harness, eval, and reward design. Synchronous or async RL. How much weight staleness the algorithm tolerates. Every one of those is a place to stall, and none of them is the actual problem you set out to solve. The problem was the thing you wanted the model to do. The rest is tax. That's the difference between a product and a partner. We bring the frontier lab-grade infrastructure and meet you wherever you are. Start on Managed Training, bring your data or evals, we run the loop. Want more control, use the Training API on Serverless and iterate on your reward or recipe fast. Ready to scale, move to Dedicated for full-parameter runs. Want people in the room, Fireworks Lab embeds researchers, engineers, and a PM with your team, from co-design all the way to a custom build we hand over. Same platform the whole way up, every model already enabled for training on elastic compute. You're not locked to one model, one GPU type, or one cluster size. The pushback I sometimes get is "we'll just do this in-house," and my answer is "at what cost?" How much do you value time to market? And where do you draw the line on quality? Customers tell us they get 2-4x more iteration on the same budget with us than with DIY or other setups. For most, the higher-ROI move is to put their people on data, evals, and product taste and let us do the rest better, faster, and cheaper. Because none of it matters if the model isn't good enough to ship. That's the part people miss. Done right, a specialized model doesn't just keep up with the frontier on your task. It beats it. One more thing on cost, because people benchmark it wrong. It's not about cheap GPU-hours. It's quality per GPU and how fast you iterate on performance/quality. We've had large tech teams run full-parameter RL on Fireworks with a fraction of the GPUs their in-house setup needed. Cheap hours don't help if you need twice as many and move half as fast. We've run RL on 10,000+ GPUs, so this holds at real scale. Full-parameter, not just LoRA. Renting intelligence gets you the average of everyone's tasks, at a premium, on someone else's roadmap. Owning yours is a bet you should make with the right partner. Build your own frontier →
もっと見る