Register and share your invite link to earn from video plays and referrals.

Benny (Yufei) Chen
@the_bunny_chen
Co-founder
101 Following    774 Followers
Specialized intelligence is hot in case you didn’t notice: @harvey’s Tenet Cognition’s SWE Cursor’s Composer GenSpark’s DeepResearch … the list goes on Join the movement with @FireworksAI_HQ training platform
Show more
Introducing Tenet, our first model post-trained for legal. Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work. Training increases Tenet's all-pass rate by 82% on LAB and 22% on LAB Contracts relative to the Kimi K3 base model. It achieves state-of-the-art performance on LAB Contracts and places second on LAB. These gains generalize to other leading agentic benchmarks including @mercor's Apex Agents - Corporate Law, @crosbylegal's Redline Bench, and @scale_AI's Professional Reasoning Bench. Tenet is also optimized for token efficiency, operating at less than a fourth the cost of leading foundation models. We additionally post-trained three specialist models for Tenet to use as subagents: 1) M&A Diligence: post-trained with @baseten on our LAB Diligence environment in an RLM harness, this model is optimized for high-scale, long-horizon tasks. 2) Review Tables: trained with @appliedcompute on our Review Table environment, this model is state-of-the-art and cost-effective at high-volume document review and structured data extraction. 3) Firm Knowledge: trained with @EngramLab on our synthetic law firm environment, this model is optimized to learn and search over a firm's knowledge via memory and structured notes. More details on model training, environment design, benchmarking, results, and more in the article by @gabepereyra below. What's next for Harvey’s research? - Scaling LAB to more jurisdictions, practice areas and workflows - Scaling compute to bring new generalist models and capabilities to Harvey More to come soon.
Show more
Introducing Tenet, our first model post-trained for legal. Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work. Training increases Tenet's all-pass rate by 82% on LAB and 22% on LAB Contracts relative to the Kimi K3 base model. It achieves state-of-the-art performance on LAB Contracts and places second on LAB. These gains generalize to other leading agentic benchmarks including @mercor's Apex Agents - Corporate Law, @crosbylegal's Redline Bench, and @scale_AI's Professional Reasoning Bench. Tenet is also optimized for token efficiency, operating at less than a fourth the cost of leading foundation models. We additionally post-trained three specialist models for Tenet to use as subagents: 1) M&A Diligence: post-trained with @baseten on our LAB Diligence environment in an RLM harness, this model is optimized for high-scale, long-horizon tasks. 2) Review Tables: trained with @appliedcompute on our Review Table environment, this model is state-of-the-art and cost-effective at high-volume document review and structured data extraction. 3) Firm Knowledge: trained with @EngramLab on our synthetic law firm environment, this model is optimized to learn and search over a firm's knowledge via memory and structured notes. More details on model training, environment design, benchmarking, results, and more in the article by @gabepereyra below. What's next for Harvey’s research? - Scaling LAB to more jurisdictions, practice areas and workflows - Scaling compute to bring new generalist models and capabilities to Harvey More to come soon.
Show more
0
122
3.1K
285
Forward to community
So Qwen3.8 27B is actually SOTA for computer use 😱😱
When should you start post-training your own models? @FireworksAI_HQ CEO @lqiao’s answer: after product-market fit. Not because it's hard... but because only after PMF is the data coming off your product surface worth training on. Lin joined us for our @sequoia "Own Your Intelligence" event to host a workshop on all things post-training; what works, what breaks, and how not to let the model outsmart you. Must listen!! 00:00 Introduction 00:37 What Fireworks sees across thousands of AI applications 02:47 Off-the-shelf APIs and the problem of keeping your taste 03:58 What "owning your intelligence" actually means 05:43 The progression: prompting → RAG → SFT → preferences → RL 07:20 Why this mirrors how humans learn 09:03 Matching the technique to the problem you actually have 10:46 Where teams get stuck: data quality and vibe evals 12:28 Reward hacking: the model that wrote zero lines of code 13:59 Training-to-serving alignment (and why quality drops) 15:55 Post-training in healthcare and security 17:31 From coding to every co-work domain 19:35 Incumbents, cost burden, and not scaling into bankruptcy 21:26 How much control do you want? 23:24 Q&A: What makes a good reward signal 25:00 Q&A: When to start thinking about post-training
Show more
OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance for a much cheaper price. Now, it is increasingly an existential and strategic topic for our portfolio. Intelligence is the product. Companies want to shape it and own it and let it compound within their own walls. Not your weights, not your product. Now, with frontier open-weight models and fantastic tooling/infrastructure, owning your intelligence at the frontier is finally becoming possible. The result: every application company we work with is embarking on the journey of doing their own research on post-training, evals, harnesses, etc. The hottest neolabs may just be @Harvey, @FactoryAI, @Ramp, etc. The list goes on. We held a summit @sequoia to convene our portfolio on this topic, together with @gabepereyra (@Harvey) on building Harvey Labs, @lqiao (@FireworksAI_HQ) on post-training, @hwchase17 (@LangChain) on harnesses + evals, @BrendanFoody (@mercor) on RL environments and synthetic data, @QuantumArjun (@trajectorylabs) on online continual learning. Opening talk below; rest to come this week! 00:00 What is sovereign AI (and what it isn't) 01:24 Centralized vs. decentralized intelligence 02:54 Four reasons companies own their models: cost, speed, performance, destiny 04:22 "Not your weights, not your product" 05:32 The application companies are the newest neo labs 07:05 Step 1: Deciding what to own vs. rent 09:51 Step 2: Build the team (and don't shoehorn your platform team) 11:17 Step 3: Legibility – why your research has to be visible 12:33 Step 4: The technical roadmap 13:56 The stack: production vs. development 15:16 Opening Pandora's box – base models, harnesses, context
Show more
Today we're announcing dfs-large1, our newest cybersecurity model that achieves best-in-class performance on vulnerability detection tasks. Besides frontier AI labs, only a handful of companies have built specialized models that reach the state of the art in their domain. We're proud to be the first to do it for cybersecurity. dfs-large1 is built on GLM-5.2 and post-trained with reinforcement learning inside depthfirst's security infrastructure. We evaluated it on depthfirst-bench, our benchmark of long-horizon vulnerability discovery across complex repositories, where it achieves best-in-class performance. Training improvements have not plateaued yet and we expect additional performance gains as we continue training. A huge thank you to @FireworksAI_HQ for being an outstanding training partner. Their infrastructure enabled us run large-scale reinforcement learning efficiently and iterate much faster. dfs-large1 is now in preview within the @depthfirstlabs platform
Show more
At Fireworks, we believe open models and ecosystem let intelligence compound where value is created - inside every company serving a special purpose. We signed the open letter to support open weights.
Show more
Every company must own its intelligence. We've raised $1.5B Series D at a $17.5B valuation. We’ve surpassed $1B ARR and serve over 40 trillion tokens daily, with 95%+ coming from models specialized on customer data. We’re just getting started. More:
Show more
Anthropic is at 5Q tokens/month Fireworks is at 1Q (>30T/day) 20% of volume with good room to grow as intelligence becomes specialized to use case btw, Q=quintillion, crazy numbers
🎙️ We welcome Benny Chen @the_bunny_chen, co-founder of @FireworksAI_HQ, to the show to explore what actually makes an AI application good or not, how to balance qualitative signals with quantitative metrics when evaluating AI, and how open-source eval protocols and community efforts are setting the standard for AI evaluation.
Show more
We are proud to formally announce Mirendil. Our mission is to accelerate science and technology through democratizing frontier AI R&D. We’re a small team with a singular focus. Our founding team consists of 20 researchers and engineers from frontier institutions including Anthropic, xAI, Google DeepMind, and OpenAI, united by a passion for science and a drive to build the technologies that move it faster. If you want to build the system that builds systems, join us! We’re fortunate to work with @a16z and @kleinerperkins, who led our seed round of $200M, followed by a major investment from NVIDIA, among others.
Show more
RT @FireworksAI_HQ: The Boardroom Club and @YossiGarmazi hosted @the_bunny_chen to talk open source, RFT, and why the infrastructure layer…
Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits, while GLM-5.2 (high) strikes a strong balance between performance and token efficiency - MIT-licensed open weights - Same API pricing as GLM-5.1 Tech Blog: Weights: API: Coding Plan: Chat:
Show more
0
726
13.2K
1.8K
Forward to community
Moonshot released K2.7 Code, the latest in their K2 line of coding models, and it's live on Fireworks Day 0, on serverless and the API. It produces roughly 30% fewer reasoning tokens than K2.6 while scoring higher on Moonshot’s coding benchmarks. For agentic coding work, that difference adds up fast. Here’s how ↓
Show more
Why I think Anthropic's uneven safety policies with the release of Claude Fable 5 undermine the broader AI community's cohesion and accelerate us to more uncertainty and risk in AI's near-term evolution.
Show more
Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
Show more
0
237
3.6K
480
Forward to community