Register and share your invite link to earn from video plays and referrals.

Eugene Ye
@gpugene
ops & compute @wafer_ai
348 Following    1.6K Followers
The biggest bottleneck in the GPU industry at the moment is not chips, CoWoS, or even power. It's credit. Sharing some thoughts on compute markets, residual value, and where compute financing is heading. Substack link in the comments.
Show more
The biggest bottleneck in the GPU industry at the moment is not chips, CoWoS, or even power. It's credit. Sharing some thoughts on compute markets, residual value, and where compute financing is heading. Substack link in the comments.
Show more
The biggest bottleneck in the GPU industry at the moment is not chips, CoWoS, or even power. It's credit. Sharing some thoughts on compute markets, residual value, and where compute financing is heading. Substack link in the comments.
Show more
claude fable 5 (left) vs gpt-5.5 (right)! "makes a twitter clone in minecraft" holy anthropic mog
0
28
2.9K
51
Forward to community
hiring exceptional engineers to join us at wafer. our next mts will work across kernels, engines + serving and own customer workloads through prod $200k-$300k + equity sf, in person 5 days/week link in comments
Show more
cpu lowkey more important in inference than gpu
wafer is growing exponentially and there's orders of magnitude more customer demand than we can attend to. we're 8 people supporting millions of ARR with triple digit month over month growth. we just announced our $40m series A to build the team that can capture that demand and make Wafer the leading inference company in the world. four roles that matter the most right now: member of technical staff work across kernels, inference engines, serving infrastructure, and heterogeneous clusters. most of your time is engineering, and you own the customer workload from benchmark through production. $250k base + plus equity. founding growth build Wafer’s marketing function with the founders and gtm lead. own positioning, founder-led social, launches, events, customer stories, the website, and distribution. $140k–$180k. chief of staff work with the ceo on whatever matters most across customers, recruiting, finance, legal, operations, and strategy. this role is for someone who can take an incomplete thought and return with a finished outcome. gtm become our first dedicated top-of-funnel hire. create qualified meetings with ai-native companies, turn them into benchmarks, and grow into full-cycle technical sales. $100k base, $140k ote, uncapped commission, plus equity. all four roles are full-time and on-site in San Francisco, five days a week. if you've ever done anything exceptional, reach out. we'd love to chat. links in comments.
Show more
RT @wafer_ai: we're hosting game night this saturday!! this week's theme: mafia dm @gpugene for an invite :)
hiring for mts, growth, gtm, and chief of staff. links in thread
cuties
Wafer (@wafer_ai) is building a fast AI inference cloud that uses agents to optimize GPUs and run open-source models at industry-leading speeds. The result is the low cost of open source with the latency of a much smaller model. Four months after launching the cloud, they went from zero to $8 million in ARR and just raised a $40M Series A. In this episode of Founder Firesides, YC's @snowmaker sits down with Wafer co-founders @gpusteve and @gpuemi to talk about going from UChicago undergrads to building one of the fastest-growing AI infrastructure companies in just over a year. They share how a college side project became Wafer, how they got GLM 5.2 running 2-3x faster than anyone else in the market, and why speed, not just cost, is pushing enterprises toward open-source models. 00:48 — What Wafer Does 01:35 — Going From $0 to $8M ARR in Four Months 04:21 — Scaling Faster Than They Could Keep Up 09:14 — How AI Makes AI Faster 11:21 — The Pivot That Changed Everything 14:41 — Why Open-Source Models Are Taking Off 18:03 — From a College Side Project to Wafer 23:33 — Advice for Starting and Joining Startups
Show more
Leading Wafer's Series A tl;dr @MarathonMP is co-leading Wafer’s $40M Series A with our friends at @Chemistry, with participation from @AMD Ventures, @Wing_VC, @outsetcap, @fiftyyears, @ycombinator, and several amazing angels @JeffDean (CEO, @DiscoLoopAI), @rauchg (CEO @vercel), @andyfang (CTO, @DoorDash), @akothari (COO, @NotionHQ), @kvogt (CEO, Bot), @eastdakota (CEO, @Cloudflare) and deepgramscott (CEO, @DeepgramAI), among others. Everything will be inference. Inference is and will be the single largest market within AI. It's the white hot center within the white hot center of AI. Today, to make inference work well, highly specialized engineers manually tune models, serving engines, kernels, and hardware. The work is slow, expensive, and often repeated when traffic changes, models improve, or new chips arrive. Wafer has built a fundamentally difference inference architecture from the ground up. Wafer's agents continuously study production traffic, identify bottlenecks, test changes, verify the results, and deploy the best-performing configurations. Simply put, Wafer builds AI that improves AI. And boy, does it improve AI. Customer after customer we spoke with, raved about Wafer's performance compared to incumbent inference providers. A very large platform said that they plan to move 100% of their inference volume to Wafer as soon as possible. Wafer cofounders @gpuemi and @gpusteve have a deep understanding of this problem through working on GPU and kernel optimization for an extended period of time. One of the most impressive things about Wafer - besides their product - is the exceptional team that Emilio and Steven have built. They have built a distinct, compelling, talent-dense culture and created an incredibly high bar for hiring that i've observed in every generational company I've been part of. Wafer has already doubled in size in the weeks since we invested, and continues growing at a rate that puts them in rarefied territory. More importantly, their product is unique, differentiated and we continue to get unsolicited raves from their customers. We’re incredibly excited to support the @wafer_ai team as they pursue their mission to maximize intelligence per watt and build a generational company that will matter for decades. PS: h/t my partner @ChaseAPackard for the meme below.
Show more
we're also hiring for MTS, CoS, Growth, and GTM dm if interested
i'm sad to announce that our recent fundraising round did not go according to plan... we initially planned to raise $18m. we didn't end up getting that number. we ended up raising a $40m series a instead. co-led by @MarathonMP and @chemistry, with participation from @Wing_VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer_ai. we are also joined by an incredible list of angels, including @JeffDean (CEO, @DiscoLoopAI), @rauchg (CEO, @vercel), @andyfang (CTO, @DoorDash), @kvogt (CEO, Bot), @akothari (COO, @NotionHQ), @eastdakota (CEO, @Cloudflare), @deepgramscott (CEO, @DeepgramAI), and more. back to the kernel mines
Show more
i get to spend all of this if you have compute please reach out Co-led by @MarathonMP and @chemistry, with participation from @Wing_VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer_ai. we are also joined by an incredible list of angels, including @JeffDean (CEO, @DiscoLoopAI), @rauchg (CEO, @vercel), @andyfang (CTO, @DoorDash), @kvogt (CEO, Bot), @akothari (COO, @NotionHQ), @eastdakota (CEO, @Cloudflare), @deepgramscott (CEO, @DeepgramAI), and more!
Show more
i'm sad to announce that our recent fundraising round did not go according to plan... we initially planned to raise $18m. we didn't end up getting that number. we ended up raising a $40m series a instead. co-led by @MarathonMP and @chemistry, with participation from @Wing_VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer_ai. we are also joined by an incredible list of angels, including @JeffDean (CEO, @DiscoLoopAI), @rauchg (CEO, @vercel), @andyfang (CTO, @DoorDash), @kvogt (CEO, Bot), @akothari (COO, @NotionHQ), @eastdakota (CEO, @Cloudflare), @deepgramscott (CEO, @DeepgramAI), and more. back to the kernel mines
Show more
0
257
1.3K
56
Forward to community
sometimes i be facetiming myself. just to hear what a real one gotta say.
wonder how many traces referenced the matrix
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Show more
We just released a massive update on our gpu performance engineering resource list AI Performance Engineering v2 this is the most comprehensive resource list for learning gpu and ai perf engineering link in thread 🧵 this version starts with how a single inference request works, then builds through the cuda execution model, roofline, transformer arithmetic, ttft/tpot/goodput, and kernel optimization. it's an opinionated list on what we think is important to deeply understand how to optimize inference systems. also added: - flashattention-4, blackwell tensor memory, and low-precision tensor cores - continuous batching, kv-cache systems, quantization, speculative decoding, and structured decoding - moe serving, collectives, topology, and prefill/decode disaggregation - blackwell ultra, mi350/cdna 4, ironwood, and trainium3 - kernelbench-verified and sol-execbench, plus a separate watchlist for rubin and cdna 5
Show more
Compute Strategy for VCs: AI-focused VCs raising funds now should 1.5-2x the amount they are targeting and spend the additional on reserved compute facilities that they allocate to portcos. Forward-thinking VCs are already doing this for late-stage funds and this strategy is starting to be adopted by growth and early-stage funds. Founders: If your lead VC cannot find you compute, that's not a lead, that's a co-investor. LPs: If a VC says they are investing in AI startups, ask them their compute strategy. 1. AI portcos will spend most of their money on compute anyway, so it is more efficient for VCs to directly acquire compute and benefit from economies of scale (pricing power, hard-to-acquire tacit knowledge) instead of each portco wasting time on it. 2. Compute will be very hard to find next year, so having compute becomes a serious capital differentiator i.e. access advantage, at the same time that most venture capital is becoming commoditized due to e.g. SPVs, crossovers. 3. Compute will likely become more expensive, so reserving compute now and offering "$50M worth of compute" in 6, 12 months may actually cost VCs much less than that. 4. VCs are positioned to bridge the creditworthiness gap between datacenters (i.e. their financiers) and startups. Because startups are new and unknown, they have to pay worse economic terms for GPUs and are often pushed to the back of the line for strategic reasons. The way that most startups get around this is to raise more money and lean on the prestige of their VCs. Consequently, VCs can and should make compute cheaper for their startups by loaning their prestige and name to compute procurement, directly. Example fund math and how to allocate a cluster: A serious early-stage fund should be raising $100-$200M today (If you want to do Seed and A in AI with less than that you are kidding yourself). Out of a $200M early-stage fund, perhaps 40% should be reserved for initial capital, 30% for follow-on, and 30% for compute (I'm eliding fees and overhead). That gets you $60M for compute. At a fictive price of $4.8 / GPU-hour for B300s, that's just about enough for a 64 node B300 cluster (512 GPUs) for 3 years. With a 64 node / 512 GPU cluster you should plan to allocate roughly like this (illustrative): * 1x 32 node block dedicated to large training jobs, held for 3-6month chunks * 1x 16 node block as a bridge cluster for R&D / training, held for 3 month chunks * 3x 4 node blocks as bridge clusters for inference, held for 1 month chunks * 4x 1 node blocks as flex blocks for compute grants to potential founders, nonprofit support, "spare change under the couch cushion" type capacity for portcos, held for 6-8 week chunks. Depending on your portfolio needs, you will want to adjust this schedule, but you want to hold a very large subset of the GPUs aside for large training jobs and as bridge capacity, a small amount as flex capacity and grants, and a medium amount for inference to help bridge for portcos. To be blunt, this is a very small amount of compute for neolabs, and only helps with bridging shortages, not with long term needs or ramps. But it hopefully provides a sense of the minimum sizes that will be needed to play in this space. That's one example. More aggressive VCs will reserve larger clusters and count on follow-on funds to pay the remaining term, or explore alternative vehicles or liquidity lines to pay for the largest cluster they can get their hands on. Finding, pricing, and closing on compute is challenging today and will get a lot harder in Q1. I'll have a longer blog post about this soon.
Show more