Register and share your invite link to earn from video plays and referrals.

Lin Qiao
@lqiao
Cofounder and CEO of @FireworksAI_HQ
249 Following    225K Followers
Excited for open model drops in Sept. 2 big ones. Very balanced.
Real cost matters! Per token is meaningless. Real cost goes beyond caching. API verbosity and accuracy matter too. An API 2x more verbose is 2x more tokens and cost. An API less accurate, compounding over hundreds of turns, is costly for end users to keep asking agents to tweak and redo work.
Show more
GLM-5.3-Flash is live on Fireworks. Day 2. Why not Day 0? Because being first isn't the goal. Being correct is. We unplugged peculiar behaviors of over-thinking from the initial tests, and shared all fixes back to open source. Quality trumps hype. Let's hold a high quality bar together.
Show more
Excited about Ox Alpha, knowing which model it is. Will get it on Fireworks.
0
94
1.4K
22
Forward to community
We are excited to drive the research work of Tenet with Harvey, delivering frontier quality across 24 areas of corporate law. Tenet exceeded Opus5 and Fable at many dimensions, covering 1300+ legal tasks. There are many good findings from this work. Congrats @harvey team for launching Tenet!
Show more
Running fast is not enough, you need fast AND correct An excellent addition from @ArtificialAnlys to make sure that the flashy speed numbers are backed by 100% matching accuracy
Show more
Announcing the Artificial Analysis Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves. We are initiating coverage with GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 coming soon Providers trade off accuracy to optimize for speed and cost. They quantize weights, write custom kernels and tune their inference stacks, and sometimes they simply ship bugs. We are bringing the rigor of our Artificial Analysis Intelligence Index to measuring endpoints, so developers can pick providers on accuracy, not just price and speed We benchmark each serverless endpoint against our own self-hosted reference deployment of the official weights, where 100% represents matching the reference. An endpoint is at reference parity when its result falls within the 95% confidence interval of the reference. Coverage is live for GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 accuracy coverage launching soon Key elements of the Endpoint Accuracy Index: ➤ Three areas, equally weighted: tool calling (BFCL-500, 500 questions, 3 repeats), scientific reasoning (HLE-250, 250 questions, 10 repeats) and long context recall (AA-LCR-25, 25 questions, 10 repeats). Each subset separates endpoints on the serving choices that drive accuracy differences, with repeats sized for tight confidence intervals ➤ Reference deployment: we self-host the official weights at the lab's recommended precision, following the lab's serving recipe, and publish the complete commands for each reference ➤ Inference parameters: we run the model's highest supported reasoning mode and each endpoint's highest supported output length and context window ➤ Confidence intervals: the parity test accounts for uncertainty in both the endpoint's runs and the reference's runs ➤ Rotating coverage: models enter once sufficient number of providers serve them and exit when a newer version in the same family supersedes them. We benchmark new endpoints as providers launch them and refresh all listed endpoints periodically ➤ Point in time: each result carries the date it was measured, with multi-day benchmarks dated to their final day Key results for GLM-5.2 ➤ Output token limits restrict accuracy. Restrictive limits cut responses off before the model finishes reasoning, and the most restrictive endpoints score half the reference or less on HLE-250 Key results for gpt-oss-120b ➤ Tool call handling separates endpoints. Providers parse and format tool calls differently, and some endpoints score 22% on BFCL-500 against 37% for the reference ➤ Serving configuration changes what the model does at the same requested settings. Some endpoints produce far fewer reasoning tokens at the same configured level, and restricted context windows truncate long context tasks Key results for DeepSeek V4 Pro ➤ DeepSeek V4 Pro endpoints are more in line with the reference. Majority of the endpoints are at reference parity, and DeepSeek's own first-party endpoint scores slightly above the reference
Show more
Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines don’t eliminate pathogens. They teach the immune system to recognize, adapt, and respond faster as threats evolve. Attackers will keep mutating. AI will accelerate that. Only if we can simulate what attacks look like, we can defend. Open models are the best tool to drive both offense simulation and defense. We as a community should build an adaptive immune systems for software.
Show more
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
Show more
We are going to see a lot of vertically focused AI native companies accelerate. Routers, open-source models and specialized post-training enabled by companies like @FireworksAI_HQ have all made dramatic advances and the combination of the three is driving accelerating growth. Companies like @wearelegora can now use their data to post-train an open-source model and then combine it with frontier models behind a router to get the same or better outcomes at lower costs than the frontier alone. This dramatically improves the business model for all these companies. @cognition seeing similar trends.
Show more
0
88
1.5K
139
Forward to community
Moonshot AI has collated a list of the Kimi K3 vendors with comparison to official API. Only @modal and @FireworksAI_HQ have submitted full results. Seems like Fireworks is the closest?
Show more
Excited to be on the CNBC live show!
Back from vacation and LIVE at 12pm PT / 3pm ET Is AI’s easy-money era ending? We’ll unpack a wild week for the AI trade—big tech earnings, Leopold Aschenbrenner’s massive deal, OpenAI’s price cuts, rogue agents, open models and new restrictions on Chinese robot imports. Cisco’s @jpatel41, Fireworks AI's @lqiao and Standard Bots' @evanbeard.
Show more
Open weights are a defender's advantage. dfs-large1 from @depthfirstlabs matches frontier-model performance on vulnerability discovery and was built on the open GLM-5.2 model, post-trained with RL on @FireworksAI_HQ. This is exactly the case for open-source AI policymakers need to notice: American companies can build best-in-class domain models using open weights.
Show more
Today we're announcing dfs-large1, our newest cybersecurity model that achieves best-in-class performance on vulnerability detection tasks. Besides frontier AI labs, only a handful of companies have built specialized models that reach the state of the art in their domain. We're proud to be the first to do it for cybersecurity. dfs-large1 is built on GLM-5.2 and post-trained with reinforcement learning inside depthfirst's security infrastructure. We evaluated it on depthfirst-bench, our benchmark of long-horizon vulnerability discovery across complex repositories, where it achieves best-in-class performance. Training improvements have not plateaued yet and we expect additional performance gains as we continue training. A huge thank you to @FireworksAI_HQ for being an outstanding training partner. Their infrastructure enabled us run large-scale reinforcement learning efficiently and iterate much faster. dfs-large1 is now in preview within the @depthfirstlabs platform
Show more
You can now fine-tune Kimi K3 on Fireworks. Conduct supervised fine-tuning, preference tuning, and reinforcement learning via Training API. Run across dedicated, and serverless training. The first open frontier model at 3 trillion parameters, ready for your product. Contact us to get started:
Show more
After signing the open letter in support of open weights Fireworks AI President @GeorgeHuSF joined us to explain why: "We don't think intelligence should be controlled by just a few frontier labs." "It should be owned by every company around the world." "The more open it is, the more access everyone has to intelligence." "We think the better society and the world is going to be as a whole."
Show more
What a week with Kimi K3 and Opus 5 launch! We compared these two great models across SWE (480), Algorithmic(100), Terminal (83), the task level quality is very close with Opus 5 being on par or better, and K3's per-task cost is 2x to 4.6x cheaper, using serverless pricing. Key measurement is per-task cost, not per-token cost. In general, open models tend to be more verbose than close models. Hence the quick study. We aim to continue to increase per-task serving efficiency (via Fireworks Inference) and quality (via Fireworks Training). Share what you find out for the tasks you care about. 👇 More details of our study --
Show more
At Fireworks, we believe open models and ecosystem let intelligence compound where value is created - inside every company serving a special purpose. We signed the open letter to support open weights.
Show more
$1BN annualized runrate 200 employees 3.5 years Such fun to sit down with @HarryStebbings and discuss the future of open/closed and specialised vs generalised intelligence. The future isn’t in duopoly. It’s in millions of companies building special products. They must own their intelligence.
Show more
Everyone gets angry with me for saying triple, triple, double, double is dead. Fine, I do not really care. Venture is about investing in unbelievable outliers. Anomalies that own markets with generational founders. @FireworksAI_HQ is an example of this. They scaled to $1BN in ARR in just 3.5 years. They also have just 200 employees making it an insane $5M revenue per head. @lqiao just raised a whopping $1.5BN at a $17BN valuation and I sat down with her. Added my notes and the episode below: 1. The Challenges That Come From Such a Fast Development Cycle for Chips Hardware innovation is moving so quickly that rapid SKU cycles now outpace traditional depreciation timelines, changing the financial calculus of building versus renting infrastructure. Founders should prioritize growth and market agility over immediate gross margins, avoiding premature optimization until customer workloads stabilize. 2. Why National Sovereignty Is Real in AI and Every Company Should Have Its Own Model Frontier models function like a society’s core electricity grid. Relying entirely on a third-party API creates the existential risk of sudden disconnection, making model ownership and infrastructure independence critical for both sovereign nations and enterprise businesses. 3. Why the Future Is Millions of Specialized Models Frontier providers bake their own design tastes and values into models, which inevitably misaligns with enterprise business logic. The future belongs to millions of specialized, “one-size-fits-one” models tailored to proprietary data, consistently outperforming generalized AGI on accuracy, speed, and unit economics. 4. The Transition From the Year of Coding to the Year of Co-Work AI adoption has rapidly evolved from engineering-centric coding tools to a diversified ecosystem of B2B co-work agents. Founders and VCs must look past crowded developer environments to capture massive value in specialized workflow automation across legal, finance, healthcare, and other enterprise functions. 5. How a 10x Cost Reduction Will Drive a 100x Explosion in Usage Temporary supply chain backlogs will eventually ease, compressing infrastructure and model-tuning costs by 10x over the next three years. This deflation in token unit economics will turn intelligence into a near-frictionless commodity, driving a massive surge in enterprise production usage. 6. Why Avoiding the Application Layer Is Essential for Platform Focus Platform defensibility requires strict focus on multi-chip agility without creating vertical hardware or software dependencies. By refusing to move up into the application layer, infrastructure platforms avoid competing with their own ecosystem and maximize their value in specialized model orchestration. 7. Biggest Lesson From Working With Jensen Huang on Leadership Leadership in hyper-velocity markets is defined by rapid judgment, not executive privilege. Because critical information degrades as it moves through layers of corporate hierarchy, leaders must stay close to ground-level technical details to maintain execution speed and avoid flawed strategic calls.
Show more
This is exactly why we believe in customization. Quick context: Factory's original secret scanner was deterministic, so it either flagged things that weren't actually secrets (false positives) or missed secrets that didn't match its patterns (false negatives). Their solution was elegant: two small post-trained models—one to catch missed secrets and another to filter out false alarms. The result: models that outperform GPT-5.5 and Opus 4.8 on this specific task while running at a fraction of the cost and latency. This is our thesis in action: take a strong open model, post-train it for a specific production problem, and you can build something that's faster, cheaper, and better than frontier models at that job. Congrats to @FactoryAI. Proud they built Droid Shield 2.0 on @FireworksAI_HQ.
Show more
Well said @RamaswmySridhar! I believe the cost saving is actually much bigger, more like 4-5x. E.g., we have just reduced our GLM 5.2 cached token price by 2X to return efficiency gain to our users. Also just launched GLM 5.2 post training with zero-KLD. Frontier quality compounds business value. More to come.
Show more