Register and share your invite link to earn from video plays and referrals.

Dmytro Dzhulgakov
@dzhulgakov
Co-founder and CTO @FireworksAI_HQ. PyTorch core maintainer. Previously FB Ads. Ex-Pro Competitive Programmer
825 Following    7.5K Followers
hear me out: AI -> SI, but it stands for Specialized Intelligence also seems we picked a good day to launch the Fireworks Specialized Intelligence Index (SII)
President Trump renames AI: "From this point forward, all of United States documents and hopefully the world's will be changed to use the much more accurate term, 'super,' as opposed to 'artificial.'"
Show more
Jev is all the hype: people didn’t realize that many classification tasks don’t need frontier LLM What you may also not realize: a few mins and $2 is all it takes to build a specialized Jev for your own task
Show more
Specialized intelligence is hot: frontier for slides generation, at 1/17th of cost Trained with @genspark_ai’s unique data and harness in partnership with @FireworksAI_HQ Congrats to GenSpark! And check out their excellent blog for motivation and tech deep dive
Show more
Meet Gen-1 Slides, Genspark’s first model designed for knowledge work. Trained with @FireworksAI_HQ from an open-weight base, Gen-1 Slides now powers Standard mode in Genspark AI Slides, delivering frontier quality decks at roughly 1/17th of Opus 5’s price. Fewer midnight fixes. Better decks by default. Read the full blog:
Show more
it’s a good model congrats to @carlobaronio @silasalberti et al it’s been delightful to partner on swe-2 from Fireworks side
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
Show more
proud of all the reliability work the @FireworksAI_HQ team is doing
Training API is in GA: specialized intelligence with full flexibility algo: SFT, DPO, RFT, OPD, your own model: 1B to 3T efficiency: sync or async numerics: aligned training/rollout/inference pricing: start serverless (per-token, LoRA), ramp up to dedicated (scale, full-param)
Show more
The most ambitious companies are training models to outperform the frontier on the capabilities that differentiate their business. Today, we announce the general availability of our Training API and Fireworks Lab, making model specialization accessible to all organizations.
Show more
periodic reminder: cache hit rate and model token efficiency (# of turns, # reasoning tokens) matter more for cost than per-token sticker price working on more updates to ⬆️ cache hit rate on @FireworksAI_HQ
Show more
same model. same list price. 5.7x apart in real cost. glm-5.2 across 6 hosts: @FireworksAI_HQ 18% of list @sferenceai 30% @tensorx_ai 43% @nebiustf 100%, caches nothing the price sheet tells you almost nothing. cache hit rate is the price.
Show more
GLM-5.3-Flash is live on Fireworks on day… 2 Why? Because we take quality very seriously. We found a benchmark discrepancy we couldn’t explain, so we delayed the launch to investigate. Day 0 (Wed): we saw 2x longer thinking on reasoning-heavy benchmarks (AIME & GPQA) for open source engines compared with @Zai_org API. Same scores, worse token efficiency. Agentic benchmarks looked good. We decided to investigate further, as overthinking might become a quality problem if max_tokens are reached Day 1 (Thu): as other non-official providers launched, their APIs had thinking in the range of open-source engines: longer than We launched a private preview endpoint with disclaimers to a few customers and worked with them to assess quality Day 2 (Fri): the official API updates. We rerun benchmarks: reasoning is now similarly long, consistent with vllm/sglang. Rest of the benchmarks, both public and internal, check out too. We launched GLM-5.3-Flash publicly: More details below
Show more
Specialized intelligence is hot in case you didn’t notice: @harvey’s Tenet Cognition’s SWE Cursor’s Composer GenSpark’s DeepResearch … the list goes on Join the movement with @FireworksAI_HQ training platform
Show more
Introducing Tenet, our first model post-trained for legal. Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work. Training increases Tenet's all-pass rate by 82% on LAB and 22% on LAB Contracts relative to the Kimi K3 base model. It achieves state-of-the-art performance on LAB Contracts and places second on LAB. These gains generalize to other leading agentic benchmarks including @mercor's Apex Agents - Corporate Law, @crosbylegal's Redline Bench, and @scale_AI's Professional Reasoning Bench. Tenet is also optimized for token efficiency, operating at less than a fourth the cost of leading foundation models. We additionally post-trained three specialist models for Tenet to use as subagents: 1) M&A Diligence: post-trained with @baseten on our LAB Diligence environment in an RLM harness, this model is optimized for high-scale, long-horizon tasks. 2) Review Tables: trained with @appliedcompute on our Review Table environment, this model is state-of-the-art and cost-effective at high-volume document review and structured data extraction. 3) Firm Knowledge: trained with @EngramLab on our synthetic law firm environment, this model is optimized to learn and search over a firm's knowledge via memory and structured notes. More details on model training, environment design, benchmarking, results, and more in the article by @gabepereyra below. What's next for Harvey’s research? - Scaling LAB to more jurisdictions, practice areas and workflows - Scaling compute to bring new generalist models and capabilities to Harvey More to come soon.
Show more
genuinely curious why these lists from Ramp and Brex look so different
… otherwise it’s just sparkling back propagation from rewards proud to be a real post-training provider and partner with @harvey
You can only be considered a post-training provider if you work with Harvey
Jeff, Sanjay, Quoc, Oriol left Google and started a neolab 🤯 Many people in CS and AI (myself included) grew up looking up to them Pitch deck has a god-level background slide - they built foundations for distributed systems and AI fields: everything from MapReduce to seq2seq to TPUs Fun fact: Jeff has a list of “Chuck Norris facts” which is quite amusing (and totally true!) Can’t wait for these legends to push scientific research automation forward with Discovery Loop
Show more
We created a pitch deck to tell a handful of VC firms about us and what we were up to (a fun experience!). Here’s a few slides about our background and some of the things we’ve worked on from the pitch deck (it was fun putting together the list of people in our teams who have gone on to found a whole range of exciting companies). We are delighted to have selected @radicalvcfund and @khoslaventures to lead our initial funding round, along with participation from @lightspeedvp, @kleinerperkins, Doerr Capital (@johndoerr), and Alphabet (@Google). We’ll be working with them to close our seed round over the next few weeks.
Show more
Running fast is not enough, you need fast AND correct An excellent addition from @ArtificialAnlys to make sure that the flashy speed numbers are backed by 100% matching accuracy
Show more
Announcing the Artificial Analysis Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves. We are initiating coverage with GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 coming soon Providers trade off accuracy to optimize for speed and cost. They quantize weights, write custom kernels and tune their inference stacks, and sometimes they simply ship bugs. We are bringing the rigor of our Artificial Analysis Intelligence Index to measuring endpoints, so developers can pick providers on accuracy, not just price and speed We benchmark each serverless endpoint against our own self-hosted reference deployment of the official weights, where 100% represents matching the reference. An endpoint is at reference parity when its result falls within the 95% confidence interval of the reference. Coverage is live for GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 accuracy coverage launching soon Key elements of the Endpoint Accuracy Index: ➤ Three areas, equally weighted: tool calling (BFCL-500, 500 questions, 3 repeats), scientific reasoning (HLE-250, 250 questions, 10 repeats) and long context recall (AA-LCR-25, 25 questions, 10 repeats). Each subset separates endpoints on the serving choices that drive accuracy differences, with repeats sized for tight confidence intervals ➤ Reference deployment: we self-host the official weights at the lab's recommended precision, following the lab's serving recipe, and publish the complete commands for each reference ➤ Inference parameters: we run the model's highest supported reasoning mode and each endpoint's highest supported output length and context window ➤ Confidence intervals: the parity test accounts for uncertainty in both the endpoint's runs and the reference's runs ➤ Rotating coverage: models enter once sufficient number of providers serve them and exit when a newer version in the same family supersedes them. We benchmark new endpoints as providers launch them and refresh all listed endpoints periodically ➤ Point in time: each result carries the date it was measured, with multi-day benchmarks dated to their final day Key results for GLM-5.2 ➤ Output token limits restrict accuracy. Restrictive limits cut responses off before the model finishes reasoning, and the most restrictive endpoints score half the reference or less on HLE-250 Key results for gpt-oss-120b ➤ Tool call handling separates endpoints. Providers parse and format tool calls differently, and some endpoints score 22% on BFCL-500 against 37% for the reference ➤ Serving configuration changes what the model does at the same requested settings. Some endpoints produce far fewer reasoning tokens at the same configured level, and restricted context windows truncate long context tasks Key results for DeepSeek V4 Pro ➤ DeepSeek V4 Pro endpoints are more in line with the reference. Majority of the endpoints are at reference parity, and DeepSeek's own first-party endpoint scores slightly above the reference
Show more
one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infra problem as a research one: 100+ turn rollouts, async/pipeline RL for high utilization, train-sampling numerical alignment Fireworks training platform takes care of infra, so research can move fast we're proud to support the ecosystem of cyber defenders like depthfirst as AI adoption accelerates
Show more
Today we're announcing dfs-large1, our newest cybersecurity model that achieves best-in-class performance on vulnerability detection tasks. Besides frontier AI labs, only a handful of companies have built specialized models that reach the state of the art in their domain. We're proud to be the first to do it for cybersecurity. dfs-large1 is built on GLM-5.2 and post-trained with reinforcement learning inside depthfirst's security infrastructure. We evaluated it on depthfirst-bench, our benchmark of long-horizon vulnerability discovery across complex repositories, where it achieves best-in-class performance. Training improvements have not plateaued yet and we expect additional performance gains as we continue training. A huge thank you to @FireworksAI_HQ for being an outstanding training partner. Their infrastructure enabled us run large-scale reinforcement learning efficiently and iterate much faster. dfs-large1 is now in preview within the @depthfirstlabs platform
Show more
These are independently evaluated by @Kimi_Moonshot . At Fireworks we really sweat small details (numerics, prompt formatting, tool parsing, etc) to bring you the best models at their peak quality and speed Huge thanks to Kimi friends for working with the community to ensure the best quality. KVV is excellent
Show more
Moonshot AI has collated a list of the Kimi K3 vendors with comparison to official API. Only @modal and @FireworksAI_HQ have submitted full results. Seems like Fireworks is the closest?
Show more
“open source shouldn’t be banned” != “everything must be open source” but also to NVidia credit big parts of the toolchain are increasingly open source: driver, cutedsl kernels, etc
Show more
I’m so excited that @JensenHuang is a believer in open source now, looking forward to the CUDA and GPU driver open source release!
💯 if there’s anything to learn from cryptography development in the 90s or recent OAI-HF model breakout is that security through gatekeeping doesn’t work the open ecosystem strengthens defense and makes AI deployment safer
Show more
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
kimi k3 is frontier, not just open frontier friends from moonshot really cooked, congrats it’s a pleasure to partner with the incredible team k3 is coming to fireworks by july 27
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: 🔗 Tech blog:
Show more
The future is specialized intelligence: your data, your product, your model Thanks to all our customers and partners for building the future together and getting to this milestone. Onwards! We are hiring across many roles. Join us! DMs are open
Show more
Every company must own its intelligence. We've raised $1.5B Series D at a $17.5B valuation. We’ve surpassed $1B ARR and serve over 40 trillion tokens daily, with 95%+ coming from models specialized on customer data. We’re just getting started. More:
Show more
wouldn’t it be great if your morning coffee was $1 instead of $5? you can get equivalent of that for your agents with open models our friends at @gumloop just did and saw 80% lower cost at same quality and they spread the message with ☕️ watch the video, very creative!
Show more
there are certain costs people have stopped questioning so when one of them drops 80%, the first reaction is disbelief open-weight models will give everyone that reaction at scale this year excited to be partnering with @FireworksAI_HQ on this
Show more