Register and share your invite link to earn from video plays and referrals.

Gokul Rajaram
@gokulr
investor ( and builder (
693 Following    122.5K Followers
Changing the agent mid session as a hand-off is incredible. Great work @agentsky_dev team.
Matan (@matanSF), Eno (@EnoReyes) and the @FactoryAI team are the real deal. Stellar product, stellar company.
We have raised $200M at a $5B valuation to scale self-improving software development in the enterprise. @FactoryAI has grown to serve hundreds of thousands of developers at companies including RBC, Adobe, Nvidia, T-Mobile, and Palo Alto Networks. We will use this capital to accelerate our investments in research, product, and global go-to-market.
Show more
Congrats Nabeel and Fin team!
$20M Raised. 7 M&A’s completed in 12 months. Investors include Expa, Coinbase Ventures, and Tenet Fund. 825 million users impacted. You have probably used us already. We help your favorite fintechs move money to the rest of the world.
Show more
Look at that latency 😍 As proven by 25+ years of internet services, speed is a feature. For any app where there is a consumer (or agent!) on the other side waiting for the result, being even a bit faster results in higher conversion (and lower drop-off / churn) @wafer_ai is the choice for low latency inference 📈🚀
Show more
DeepSeek-V4.1-Flash is live on @OpenRouter!! pick @wafer_ai as your provider (ss taken 9.12.26)
This is incredible. Congrats @0xSigil !
Anyone can get my friend’s real SSN from Claude I'm Terrified of my data in training sets. Emails in databases. iMessages harvested. Cant trust the cloud as AI is crazy good at hacking So I built Underdog for myself: on-device AI OS. Capable & 100% local. now my friends love it
Show more
Ambition is the Bottleneck now @tarstarr (Tara Seshan) (@OpenAI , product lead for Codex and ChatGPT Work), interviewed by @lennysan (Lenny Rachitsky), Lenny's Podcast The easy work is now trivially easy and the hard work is easy, so what separates people and companies is the ambition behind what they are willing to attempt. Seshan runs Codex and ChatGPT Work at OpenAI. Her operating rules follow from that: build for models 2 to 3 months out, ship prototypes instead of documents, and treat raising other people's ambition as part of the job. 1. Steering Over Rowing. Agents do the rowing and people steer, and the steering keeps moving up a level. It used to be a line of code, then pressing tab, then a goal, and Seshan expects it to keep climbing. Picking the direction stays human, and she describes it as a positive determinism about what you want the world to look like rather than a readout from data. Her next problem is multiplayer: people at OpenAI were sending each other screenshots of their Codex threads in Slack, which is a poor way to work with agents together. 2. The Ambition Bottleneck. "Not only are we able to be more ambitious, we almost need to be more ambitious." The people Seshan sees getting the most out of AI use it to widen the set of things they can do at all, beyond automating rote tasks. The old unicorn hire was the product thinker who could also engineer and design, because that person removed the translation layers between functions, and everyone has that now. The constraint moved from what you can execute to what you can imagine, and expanding your own thinking is the hard part. 3. Raising Other People's Ambition. Seshan cites Tyler Cowen: people underrate walking up to someone and asking whether they could do the more ambitious version, or do it faster, or do it at 10x the scale. She treats that as a large part of the PM job, so when someone proposes a timeline or a v1 scope, the response is to ask whether the ceiling is higher. Her evidence is Patrick Collison's list of projects executed at unreasonable speed, all of which happened before these tools existed. If those were possible then, the count should be climbing fast now. 4. The Two To Three Month Rule. "You fail if you build for where the models are now. You fail if you build for where you think the models will be in a year. Both outcomes are equally wrong." Seshan builds for capability 2 to 3 months out, which only works if product stays tied to what research has on its roadmap. The discipline is putting model capability at the center and getting your own product constructs out of the model's way. She quotes Kevin Weil's line that this is the worst the models will ever be, and says it is absurd that it keeps being true. 5. Empirical Over Academic. At Stripe, payments rewarded rigor: you could reason through a competitor's next move, and failing to do that showed up as carelessness. Seshan found AI markets too emergent for that, so being prolific beats being theoretical, and the switch felt jarring enough that she wondered whether she was skipping her due diligence. What replaced the long reasoning doc is sharpening one hypothesis to a point, testing it, and feeding the result back in. She borrows Shishir Mehrotra's term for it, the eigenquestion: the one thing that determines whether the product works. 6. Founders, Plural. OpenAI is founders led rather than founder led, with very little top-down direction and almost no distance between a product lead and the market. Seshan expected a treasure trove of secret strategy, the way a new hire at Stripe gets handed the payments bible, and there was none. Every view about how the world should work becomes public product or public messaging quickly. She credits the Codex turnaround to that structure: people who noticed something should be better went and built it without asking. 7. Three Questions That Run Product. OpenAI operates on three internal questions. "Is this maximally accelerated?" came from Nick Turley and covers speed. "Are you mainlining it yet?" is the one Seshan and Andrew Ambrosino ask their team, and it means using the product all day every day to do your actual job. The third asks whether the team is being as ambitious as possible, which covers scope. 8. Knowledge Work Isn't Code. Coding is output-verifiable: run the tests and you can trust the answer. Seshan says knowledge work breaks that, because you cannot look at the finished deck, see 90%, and believe it without inspecting the process, the inputs, and the reasoning. So the product has to show in-progress work, citations, and chain of thought, and take the user along to the answer. It also puts the thread format itself in question, since threads were built for coding. 9. Writing As Thinking. Seshan splits writing at work in two and treats the halves oppositely. Writing as reporting covers status updates and launch plans, and she hands all of it to the model. Writing as thinking covers the brief that argues for a product or a strategy, and she never automates it: "I start myself and I end myself," using AI in the middle only to pull data or push back on her ideas. The same rule governs her meetings, where she prepares for the total time everyone will collectively spend in the room. 10. Mocks, Not Docs. A long document stopped being proof that you thought about something, because anyone can now generate a long document that proves the opposite. Seshan still writes hundreds of docs, and she writes them for herself. What she circulates is a prototype people can try, or better, an A/B result with a recommendation attached. She calls this the biggest personal change of the era. 11. The 70 Percent Doc. Take a document to 70% and let the people whose buy-in you need carry it to 100%. Seshan got the rule from an old manager and still runs it. A perfectly polished idea repels new ideas, which bounce off it, while something with rough edges invites people to work on it with you. At Stripe she used this to open new product areas, shopping a brief around and letting each person attack it before taking it to the next. 12. Product Marketing Fit. Seshan spent time at Sutter Hill to learn whether product market fit is luck or a playbook, and came away convinced it is a playbook. What surprised her most was how badly she had underrated product marketing fit, having treated PMM as glue between functions. Mike Speiser tests the narrative before the product exists: pitch 100 people, refine the positioning until the story lands, then commit to a product shape. Done well, that work can decide whether the company succeeds.
Show more
“A committee is a group of the unwilling, chosen from the unfit to do the unnecessary.” - @FelixDennis
The ARR multiple fallacy ARR multiple is not the right way to think about sub-20% growers. Simple framework: If you’re growing sub-20%, you get a 10-30x EBIDTA or FCF multiple. If you use this rubric, Miro is very fairly valued by Bending Spoons. Only 30%+ growers get the luxury of an ARR multiple. The reason is that 30% growth means you’re doubling in 3 years and so the revenue base will be 2x in 36 months, and the margin structure will look different at 2x scale. Cash flow today is a rounding error against that. So you value the trajectory, and ARR is the cleanest proxy for trajectory. At 15% growth, doubling takes 5 years. That's close enough to "never" that the market stops paying for the future and starts paying for the present. The present is earnings. If you're not producing them, you're not a growth company anymore. You're a bad value company. The mistake founders make is anchoring to the multiple they had at 40% growth and thinking the ARR multiple just compresses. It doesn't compress. It gets replaced. You cross 30% on the way down and the entire valuation framework switches out from under you. Which is why the worst place to be is 20-30% growth with no FCF. Growth investors won't pay for it. Value investors can't. You're priced by whoever is least excited,a really painful place to be.
Show more
This is the right time to start the Bending Spoons for AI companies. Start small today, figure out playbook, grow into 1-2b acquisitions of slow growing 600m ARR companies in a couple of years.
Congrats @fletchrichman and team @typedotcom - Type is the best team shared space I’ve used.
Today we're officially launching @typedotcom, the shared space for your team's best AI work. Teams need a central place to build apps, skills and automations together. Not swarms of agents. Our ambition with Type is to be the last piece of software your team will ever need. To build this, we've raised $4m from @LererHippeau, @haystackvc, @MatchstickVC, and angels from Slack, Github, and Ramp. So what does Type do? → share work from Claude chats into a collaborative space → switch harness or model any time → manage integrations with granular permissions → a central place to deploy apps and automations together → build a self-learning company brain → use your existing Claude or ChatGPT subscription Companies like @rayconglobal, @Intelligems and @trueclassictees are already running tens of thousands of successful agent jobs that forecast revenue, manage inventory, answer support tickets, and a lot more. We're blown away every day at the custom tools and automations our customers are building. What's something you haven't yet been able to solve or automate in your business? Tell us in the comments and we'll respond by building it for you in Type!
Show more
Work on your Introduction This gem from @george__mack is phenomenal advice for EVERYONE. Polishing your intro and making yourself into a "simple API" (!) is a great way to engineer luck. "Work on your introduction - This could be the least British advice I will ever give. I can hear my ancestors spinning in their graves at the thought of what I’m about to say. In British culture, we’re taught to play down everything that we do. “I just do marketing stuff”. The problem with this is that people you meet don’t understand what you do or how they can help you. When you have a clear introduction that describes what you do: “I create Super Bowl-level commercials for fintech companies on social media”, they can now realise ways they can help you: “Oh, my friend Barry is the marketing director at Amex. Let me introduce you!”. Being a great luck engineer is turning yourself into a simple API that people can connect into."
Show more
Outcome-based Marketing To sell AI to enterprises, you need to sell Outcomes For 20 years, software companies sold tools to enterprises. Today, the best and fastest growing companies sell outcomes, not tools. @pepper_content is doing this for marketing. Pepper got its start running organic growth programs for more than 250 large enterprises. The team was responsible for monthly results, so it learned the work in detail and documented every step. Pepper used that experience to build Atlas, a system of 365 agents. Atlas tracks the questions buyers ask, publishes pages designed to answer them, earns mentions across sources AI systems trust, and attributes the leads that follow. Agents handle 80% of the work. A senior marketer on the customer’s team owns the number. Pepper sells organic and AI search-sourced pipeline as a single line item. Its commercial model puts the company on the hook for the customer’s result. That accountability matters. AI-native services depend on detailed knowledge of the work: which decisions recur, where judgment is required, and which actions produce results. Pepper accumulated that knowledge by serving customers for years. It then encoded those decisions into software while keeping a person responsible for the KPI. @SinglaAnirudh started Pepper at 18. He has spent every year since learning what CMOs need and building the company around those needs. Building a services company is hard. Automating most of its work while remaining accountable for customer results is harder. Pepper is showing what an outcome-based AI company can look like. I’m excited to support them on this journey.
Show more
"The biggest bottleneck in the GPU industry at the moment is not chips, CoWoS, or even power. It's credit." our head of compute @gpugene on how uncertain GPU residual values shape lending terms and the capital required to bring new compute online. link to article in thread 🧵
Show more
Consumer AI is wide open @Bchesky (Brian Chesky), Co-Founder & CEO of @Airbnb, interviewed by @BrianSozzi (Brian Sozzi) (Power Players) Summary: Brian Chesky thinks Silicon Valley is mistaking today’s chatbots and coding agents for the finished product. Enterprise AI is crowded, while consumer AI—the products that actually change daily life—remains largely unbuilt. Airbnb is betting that its trust network, identity layer, preferences, payments, and travel marketplace can become an AI-native platform, while small elite teams use the tools to make a 17-year-old company move like a startup. [PS: I thought this was a particularly relevant podcast, given the massive number of new 11 personal AI assistants that have emerged on the scene over the past few weeks] 1. Growth can reaccelerate: Airbnb’s revenue grew 17% this quarter versus 10% last year. Chesky credits small elite teams, nearly twice as many feature releases, better quality, and aggressive AI adoption—proof that scale does not have to make growth obey gravity. 2. AI is an operating model: Airbnb is not treating AI as a feature team. Chesky hired a former leader of Meta’s Llama work as CTO, built internal AI tutors, and is pushing executives to run more of the company with AI. 3. The hard part is cultural: AI looked like a freezing pool until employees jumped in and discovered warm water. The tools can teach people how to use them; the real obstacle is a large company’s instinct to preserve familiar workflows. 4. Small teams regain leverage: Startups begin AI-native, but incumbents that overcome inertia can combine the speed of AI with the robustness, data, and distribution of a scaled company. That is the best-of-both-worlds prize Airbnb is chasing. 5. Every new category should strengthen the core: Fifty-five percent of guests who first book a hotel on Airbnb later book a home. Hotels add revenue and acquire customers for the core marketplace—the same flywheel that took Amazon from books to everything. 6. Aging users are an asset: Airbnb’s original 26-year-old customers are now 44, wealthier, and traveling with families, while new young cohorts keep arriving. The product can grow across Gen Z, Gen X, and boomers without choosing between retention and relevance. 7. Reinvention is chapter three: Chapter one was product-market fit and hypergrowth. Chapter two was surviving an 80% pandemic collapse and going public. Chapter three is turning a short-term rental marketplace into a trusted platform for traveling and living. 8. The platform starts with trust: Airbnb has more than 200 million verified identities. In an internet flooded with synthetic people and content, identity, remembered preferences, rich profiles, and payments can become more strategic than the booking interface itself. 9. Chat is not the final interface: Chesky believes agents are real but their current form is primitive. Consumer AI has barely touched voice, video, photos, or native interfaces, which is why most people’s daily lives have changed far less than the hype suggests. 10. Consumer AI is wide open: Of 175 companies in a recent YC batch, 159 were enterprise. Chesky sees the neglected opportunity in products regular people love—tools that make medicine, travel, and everyday services dramatically more accessible. 11. Starting is cheap; frontier scale is not: One person can now build an agent for almost nothing and may create a billion-dollar company. But frontier models can require thousands of GPUs, elite researchers, and more than a billion dollars. AI is democratizing creation while concentrating infrastructure. 12. Travel will move up the funnel: Chesky expects AI-native search, destination discovery, and customer service to transform travel in the next year. Chatbots may inspire the trip, but he does not expect them to own the reservation anytime soon.
Show more
Superb article by @gpugene on buying compute. Eugene works at @wafer_ai, and this depth and clarity of thinking is illustrative of how stacked and cracked the team is and why we at @MarathonMP are so excited to be an investor.
Show more
The biggest bottleneck in the GPU industry at the moment is not chips, CoWoS, or even power. It's credit. Sharing some thoughts on compute markets, residual value, and where compute financing is heading. Substack link in the comments.
Show more
we launched the most comprehensive ai performance engineering repo in the world last week now we'll be doing a deep dive on every single resource this is Wafer's AI Performance Engineering Series save this as your starting point. links in thread 🧵 Part 1: Prefill vs Decode a causal autoregressive Transformer computes next-token logits from the supplied prefix. a decoding policy selects one token and appends it before the model predicts the next. in a conventional causal decoder, each position attends only to itself and earlier positions. appending a token adds no allowed input to old positions. in evaluation mode, with the prefix, weights, mask and positional computation fixed, the old representations remain reusable. the KV cache stores earlier positions' per-layer keys and values for reuse in later forwards. prefill processes the known prompt under a causal mask and saves each layer's keys and values. the final prompt position produces logits for choosing the first output token. prompt positions can run together within a layer; the layers still depend on one another. the decoding policy selects a token. greedy decoding takes a highest-scoring token; sampling draws from a distribution. feeding the selected token back through the model creates its K/V and produces the logits for the next output. emitting a token and processing it are separate steps. to emit N ≥ 1 tokens, an ordinary loop needs one unchunked prefill and N−1 incremental forwards. emitting the final token does not require processing it through the model. the new token still runs through the layers, including its projections and MLP. its query computes new attention scores and a weighted sum over cached values. storing K/V saves repeated prefix computation; it leaves new attention work over the growing context. for fixed prompt length and model dimensions, caching changes total projection and MLP work from quadratic to linear in output length. full causal attention across generation remains quadratic. these work counts do not establish a measured latency improvement. prefill has many known token rows from one prompt. ordinary decode contributes one new row per active request, so batching it combines different requests. each request also needs its own logical KV state. more concurrent requests and longer retained contexts increase that state. choose the performance target around the workload. offline generation may prioritize completed work per dollar within a deadline. streaming chat also needs low time to first token and responsive delivery afterward. a single-user device may favor latency within its memory and compute limits. for a serving comparison, fix prompt/output lengths and concurrency. measure client time to first token, the intervals between later tokens, and total output tokens over a common measurement window. keep model-only prefill timing separate from queueing and delivery. a higher aggregate token rate alone cannot tell you whether one user's answer arrives sooner. chunked prefill and speculative generation change the execution pattern and are outside this ordinary-loop example.
Show more
DoorDash has really upleveled its AI products and infrastructure. It all starts with bringing the right leader / talent in. Other companies can take a page from their playbook. Kudos @aryanxshah and @andyfang for so strongly building DoorDash's AI research capabilities.
Show more
We’re excited to welcome @chahuja to the team! We're hiring for Applied AI and Core Research roles. DM us or learn more here: