Register and share your invite link to earn from video plays and referrals.

Fireworks
@FireworksAI_HQ
The frontier platform for training and inference on open-weights models at scale.
291 Following    31.8K Followers
Our partners @LangChain generate billions of tokens of agent traces a day, their richest signal on how real users react. Judging every one with a frontier closed model like GPT-5.5 or Opus is too costly, so they fine-tuned a Qwen base model on Fireworks that matches it at up to 100x lower cost. Build your own frontier:
Show more
DeepSeek V4 Flash now comes with native vision on Fireworks. Available now on serverless and priority tiers. Better benchmarks than the text-only 0731 version at the same low price: $0.22 input / $0.66 output per 1M. Give your workhorse model eyes. Try it today →
Show more
Train past the frontier. @tryheidi's ambient scribe drafts a clinical note the clinician has to trust enough to sign. They fine-tuned open models on Fireworks whose notes their clinicians preferred over Gemini’s in side-by-side reviews, at 3.5x lower latency.  Build your own frontier:
Show more
A secret that leaks into shipped code is a breach waiting to happen. @FactoryAI's droids write and commit code faster than any reviewer can keep up, so Droid Shield has to be trustworthy and low-friction: catch the real secrets, clear the false alarms. With the Training API, Factory fine-tuned an open Qwen model that caught almost 20% more real secrets than GPT-5.5 at lower cost and latency. Build your own frontier:
Show more
I am the product lead for the training platform at @FireworksAI_HQ . Today the Training API and Fireworks Lab are GA. I won’t walk you through all the features. I want to tell you why teams train with us, because it's usually not the reason you'd expect. Some teams treat training like a project. Pick a model, tune it, ship it. Done. The teams that win don't think that way. For them it's a long-term bet on owning their AI. And a bet like that needs a partner across the whole development cycle, not one piece of it. The silent killer is getting the infrastructure wrong. In RL at scale, training and inference stop being two systems and become one loop. The challenges start with model choice, GPU procurement, and room to scale when the run grows. They get harder once the loop is running. Numerics drift between the trainer and the rollout engine, and your learning signal corrupts while every dashboard says the run is fine. GPUs sit at synchronization barriers waiting on a handful of long completions. Weight sync eats the step time you thought was going to training. Most teams find this weeks in, after stitching compute, inference, and training together themselves. The second gap is expertise, and few organizations hold frontier-lab talent at every layer of the stack. Harness, eval, and reward design. Synchronous or async RL. How much weight staleness the algorithm tolerates. Every one of those is a place to stall, and none of them is the actual problem you set out to solve. The problem was the thing you wanted the model to do. The rest is tax. That's the difference between a product and a partner. We bring the frontier lab-grade infrastructure and meet you wherever you are. Start on Managed Training, bring your data or evals, we run the loop. Want more control, use the Training API on Serverless and iterate on your reward or recipe fast. Ready to scale, move to Dedicated for full-parameter runs. Want people in the room, Fireworks Lab embeds researchers, engineers, and a PM with your team, from co-design all the way to a custom build we hand over. Same platform the whole way up, every model already enabled for training on elastic compute. You're not locked to one model, one GPU type, or one cluster size. The pushback I sometimes get is "we'll just do this in-house," and my answer is "at what cost?" How much do you value time to market? And where do you draw the line on quality? Customers tell us they get 2-4x more iteration on the same budget with us than with DIY or other setups. For most, the higher-ROI move is to put their people on data, evals, and product taste and let us do the rest better, faster, and cheaper. Because none of it matters if the model isn't good enough to ship. That's the part people miss. Done right, a specialized model doesn't just keep up with the frontier on your task. It beats it. One more thing on cost, because people benchmark it wrong. It's not about cheap GPU-hours. It's quality per GPU and how fast you iterate on performance/quality. We've had large tech teams run full-parameter RL on Fireworks with a fraction of the GPUs their in-house setup needed. Cheap hours don't help if you need twice as many and move half as fast. We've run RL on 10,000+ GPUs, so this holds at real scale. Full-parameter, not just LoRA. Renting intelligence gets you the average of everyone's tasks, at a premium, on someone else's roadmap. Owning yours is a bet you should make with the right partner. Build your own frontier →
Show more
All week we’re spotlighting teams training past the frontier on Fireworks. @Figma’s AI team is shipping features that change how millions of people design, and with the Training API they can iterate faster. Less time debugging infra, more time on research. Build your own frontier:
Show more
GLM-5.3-Flash: GA on Fireworks. We held it two days over a reasoning-length gap we couldn't explain. Same weights ≠ same model. Quality first. Start building →
Show more
GLM 5.3 is now live on Fireworks, day zero: → 50% improvement over GLM 5.2 on Code Bench → State-of-the-art open-weight performance on cybersecurity tasks → Excels at complex coding and long-horizon tasks → US-Hosted Serverless on Fireworks Try it now:
Show more
Recently @Harvey introduced Tenet, its first model, trained for long horizon legal work. Harvey post-trained it from a Kimi K3 base in collaboration with Fireworks using async RL on our Training API. A thread on promising initial results for both performance and cost-efficiency🧵
Show more
Sharing the first round of speakers at the 10th Open Source AI Summit SF! 💘 @matthew_d_white: Former global CTO of AI of Linux Foundation and CTO of PyTorch @dzhulgakov: Co-Founder of @FireworksAI_HQ @lukaszkaiser: Member of Technical Staff at OpenAI @ilblackdragon: Co-Founder at @NEARProtocol Registrations are rapidly increasing - excited to meet with the community! Secure your spot here: See you next Friday at @thehousesf <3
Show more
Here's a quick look at the live cost monitoring built into Fireworks Nexus, what it tells you, and how to use it. Learn more about Nexus:
Join us at Glean:GO 2026 at Fort Mason, SF. We'll have two sessions: - Emerging open standards for the AI ecosystem, w/ @aaronamelgar - AI-Native Showcase keynote, w/ Co-Founder & CTO @dzhulgakov Did we mention we'll serve coffee? Register:
Show more
Specialized intelligence is hot in case you didn’t notice: @harvey’s Tenet Cognition’s SWE Cursor’s Composer GenSpark’s DeepResearch … the list goes on Join the movement with @FireworksAI_HQ training platform
Show more
Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous agents, heavy coding, large context windows, and is ideal for coding and agentic performance. Start building today:
Show more
You can now fine-tune Kimi K3 on Fireworks. Conduct supervised fine-tuning, preference tuning, and reinforcement learning via Training API. Run across dedicated, and serverless training. The first open frontier model at 3 trillion parameters, ready for your product. Contact us to get started:
Show more
Kimi K3 is live on Fireworks. Day 0, inference and training. US-hosted, and zero data retention. This is the first frontier open model in the 3 trillion parameter class. It sports 1M context, native vision, and reasoning that rivals the top closed models. Boom.
Show more
We ran Kimi K3 against Fable on ~1,000 agentic tasks, expecting a catch-up story. We got a specialization story instead. @kimi_moonshot's K3 outperformed on security, crypto, and long terminal loops. Fable beat on multi-lang + web/data viz. Per-task routing hits 93% accuracy, above BOTH models, at up to 50x lower cost than Fable on long loops. The part nobody's pricing in yet: the router sends 72-96% of traffic to K3. The frontier model becomes the fallback rather than the default. Kimi K3, coming to Fireworks July 27.
Show more
0
61
2.1K
181
Forward to community
Open-weight models like GLM-5.2 and Kimi K2.6 demonstrate the capability & performance needed for real GTM work. @ClayRunHQ's team recognized that early, and we're proud to be the inference layer making it real. Congrats to @jeffbarg and Clay. Let's keep building together.
Show more
Specialized intelligence just became a $17.5 billion idea. Last week @lqiao made the case on the RAISE Summit Master Stage. This week @FireworksAI_HQ announced a $1.5B Series D — $1B+ ARR, up 5x YoY, 40 trillion tokens served daily. Her argument: efficiency means specializing the tool to its value. "We don't drive a Ferrari to grocery shopping." Full session 👇
Show more