Register and share your invite link to earn from video plays and referrals.

Tanvir Bhathal
@BhathalTanvir0
cs/ai + math @Stanford, research @Google @HazyResearch, president @StanfordAIClub | prev @AIatMeta @BrainCorp
436 Following    730 Followers
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰 As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: More on verification scaling in my previous post.
Show more
0
154
3.1K
400
Forward to community
All @_wrangle search models are SOTA on people-search. 89.29%, 89.81%, and 91.10% respectively on Precision. For reference, Exa self-reports 63.3%.
We are launching Hone today. Over the last few years I had a front row seat to how AI has reshaped software engineering. Model progress exceeded my wildest expectations. Yet outside of engineering, most organizations struggle to derive value from AI. At Hone, we enable organizations to direct AI at their most complex business problems: AI that owns outcomes over weeks and months, not tasks. Read more about our mission at I am immensely grateful to @ScottWu46 , @stevenkplus1 , @walden_yan , @russelljkaplan and all the friends at @cognition for the incredible journey and everything you taught me. Thank you for your support in this new journey.
Show more
Worked with the @osmo_studio team for @stanfordaiclub, they are exceptional and have the highest bar for quality
I left @A24 to start Osmo, a storytelling studio with a software factory. Today, we’re launching @osmo_studio to the public: an end-to-end editable video platform that turns Figma, code and reference videos into fully editable motion. We also raised a $5M seed to build it.
Show more
I had so much fun creating the Self-Improving AI Agents course with @achowdhery, and teaching it twice in one year at Stanford! We also collaborated with Stanford Online to make the course available online: YouTube: The field is moving incredibly fast, but we tried to focus on the core concepts that help us build better AI systems. I hope you enjoy it!
Show more
0
40
1.1K
117
Forward to community
Must read blog from Nash & the Cursor team!
We’ve open-sourced the MoE megakernel we use to train models on NVL72s. A few of my favorite details: bf16 and mxfp8 support, pull-based dispatch, configurable overlap granularity, full determinism, and no CPU-GPU syncs. Read more in our blog and code!
Show more
(1/9) I'm thrilled to share the open-source release of Mixture-of-Kittens (MoK), our MoE megakernel for NVL72s! MoK fuses all mixture-of-experts communication and computation into a single, fully deterministic kernel, and powers Composer training across tens of thousands of GPUs. Joint work with @nash_c_brown, @hmwildermuth, @tmwilliamlin168, and @ellev3n11
Show more
@JonSaadFalcon and @Avanika15 propose measuring intelligence per watt rather than capability. It reframes the margin question. Frontier pricing does not require better models to fail. It requires only that switching costs stay near zero, lock-in stays weak, and most workloads never reach for the frontier. All three hold.
Show more
I have been talking about various similar topics for a while, Jaya frames them quite well in this
Hey @FCBarcelona any roles available?
We're hiring a Research Engineer at @Arsenal ⚽🔴⚪ to work directly with our Men's First Team! We're building state-of-the-art AI models for the football domain. This role will focus on building the application layer for our research to advance coaching and analysis workflows.
Show more
MLPs store facts in language models. Can we write them into Transformers without training? New work w/ amazing team @garctrob @ronnygjunkins @EyubogluSabri, Atri Rudra & @HazyResearch gives a ✨closed-form✨ recipe for fact-storing, Transformer-ready MLPs. Accepted at COLM 2026!
Show more
Excellent and concerning analysis on how easy it is to have false truths accepted by mass populations
“not worried about messi at all. he'll be fine. he has every trophy that matters and two decades of footage behind him. what worries me is everything else. if the internet can reshape the reputation of arguably the most documented footballer in history, imagine what it can do in situations where there isn't twenty years of footage for people to check. politics, wars, business, or people you've never even heard of. in those cases, the narrative often becomes the only version most people ever see.”
Show more
We know the J-space can surface what a model is thinking. But can it tell us how a model computes its answer? In a new blogpost, we apply J-Lens to find “meta-tokens” that can directly tell us the algorithm Qwen-3.6-27B uses to complete a task. 🧵
Show more
Assemble builds agents that help IT teams maintain enterprise systems like ERPs, CRMs, and HRIS platforms. Map the dependencies, implement configurations, and audit the agent all on one platform. Excited to launch today, learn more at
Show more
⛳️ We built a $40,000 TrackMan golf shot tracker with a Raspberry Pi and a ~$100 ti iwr6843sk radar. 🏌️‍♂️More details below!
Rumor has it the next trillionaire will be in this room
POV: you're in SF during a once-in-a-generation shift in technology. You just graduated, or you came out for the summer. AI billboards everywhere, models getting smarter every week, equal parts exciting and daunting. And somewhere this summer, you meet your future friends and cofounders. We're bringing them together for a night in SF with @sequoia. You'll be in the room with some of the founders defining this moment too. July 30th. Invite only. DM me for an invite.
Show more