Register and share your invite link to earn from video plays and referrals.

Veeral Patel
@vral
head of applied ai 💳 investing
1.3K Following    4.7K Followers
Ashwin benchmarked Jev as a replacement for LLM-based reranking in Ramp's accounting product, and the results are very promising Jev matches current accuracy on GPT-5.6 Luna, but decreases tail latency by 10x to 300ms at 3x lower cost We're excited to get this into production for 70k customers as soon as we stop running into rate limits :)
Show more
Companies are likely spending millions subsidizing employees’ personal ai use because oai and anthropic can’t figure out how to build a decent account switcher
This chart is a Rorschach test for how bullish or bearish you are on the AI trade
New from Ramp data: the latest threat to the AI trade. AI companies' revenues are heavily dependent on a small set of customers. 80% of OpenAI and Anthropic's enterprise revenues come from 1% of their customers, and it's not getting better. This is a level of concentration risk unseen in any other software category we track. The companies in the top 1% skew heavily toward the tech sector and AI products and services. What happens in a market correction? All these companies are highly correlated, and an increasing share of our economy is invested in them. Especially as we approach blockbuster IPOs for OpenAI and Anthropic.
Show more
Ramp and are going to integrate seamlessly to help your business understand what work your tokens are producing to measure ROI DM or email me if you'd like a preview
The Modern Data Stack is over, long live the Post-AI Data Stack🫡🤖 I wrote about how technological shifts change what a data team can build, the features of post-AI data stacks, what we've built at @tryramp, and what data teams can learn from @nbcsnl.
Show more
The companies of the future will allocate compute like capital: continuously routing it toward high ROI outcomes Ramp is building the tools to help
Right now, traffic is compounding at roughly 34% per weekday. It’s early, but the product is useful and the growth curve is as steep as any I’ve seen.
Fiduciary duty may not require directors to choose open-weight models, but it should make it increasingly uncomfortable for them to approve material proprietary-model spending without knowing whether open alternatives can perform the same work for less. Now that credible open-weight alternatives offer material savings, greater data control or reduced vendor concentration, that process should normally include comparative benchmarking and total-cost analysis. As AI spend becomes material and credible alternatives offer potentially order-of-magnitude savings or materially different security and dependency risks, the duty-of-care expectation of an informed process increasingly supports a documented alternatives analysis.
Show more
Ramp is a "save you time and money" company is our first step to go beyond fintech DM me if this is exciting to you- we're actively working on money saving products across the AI stack (router, harness, application layer)
Show more
the request-level evidence is the interesting part: Default baseline: $0.001219 Flex actual: $0.000610 same model, roughly half the cost. this is the first test that gave me a concrete reason to use Router instead of calling Luna directly. Nice work, @tryramp , @vral and team! next thing I’d love: aggregate baseline vs actual spend, flex share, fallback count, and why each tier was selected.
Show more
try Fable in the codex harness with Honestly it's been game changing for me.
We're grateful for the incredible response we’ve gotten since the launch yesterday. The mission to get more out of AI spend is clearly resonating, and we’re excited to be partnering with @OpenAI to continue doubling down here and make AI even higher ROI for businesses.
Show more
the frontier just got more affordable. we teamed up with @OpenAI: GPT-5.6 Sol is 50% off on router through September 18. one endpoint, every model, each request routed to the best fit. so every AI dollar goes further. start saving in a couple minutes at
Show more
Monitor and control your AI spend on every provider on Our early users save 40% on average. Every week, the price-intelligence-latency frontier shifts, and we expect this trend to continue. Tradeoffs between latency, reasoning, cost, service tier, open source and closed source models are shifting constantly. Router sends every request to the model that's actually best for the task and helps you control what tokens you buy. We benchmark it against real work: ~40% lower cost for the same outputs. Today we're opening it to everyone. Two lines of code or just change your base URL. No @tryramp account needed. Free through 2026, first $26 on us. Get an API key today at
Show more
0
147
1.4K
113
Forward to community
Everyone is underestimating NVIDIA's ability to innovate up and down the AI stack We implemented NVIDIA Switchyard on our internal SWE-Bench and cut costs by 58% It's live in Ramp Router today - we're actively removing people from the waitlist while we manage load!
Show more
We tested NVIDIA NeMo Switchyard’s stage router for coding agents in Ramp SWE-Bench. Routed agents showed comparable performance to single-model controls while substantially reducing costs and runtime.
Show more
there’s no substitute for being obsessed and making sure every token counts 🫡
There is no substitute for looking at the data. Manually inspecting failure modes is how you discover the shape of the problem. Great AI products are built by teams who set up their infra + instrumentation to make looking at data easy. @austospumanto @bcherny @rahulgs
Show more
The company that figures this out is going to save knowledge workers a lot of time and money
What’s the best open source version of Claude Cowork that a) allows you to plug in @OpenRouter or @bittensor b) works with local models and c) is optimized for the average knowledge worker (Hermes and OpenClaw too complex for normies) Right now Perplexity computer is my best choice (doesn’t do local models yet)
Show more
Set up Ramp Router in ~5 min for my vibe coded workout tracker. Tested about a dozen models on different use cases, and found pretty comparable performance between them for most basic tasks. The app is now 94% cheaper and running primarily on deepseek-v4-flash
Show more
Every week, the price-intelligence-latency frontier shifts, and we expect this trend to continue Across 100+ use cases in our product, keeping all up to date with the right model is a challenge - Either we're losing out on intelligence for the dollars we spend, or we're spending too much money for the intelligence the feature needs Ramp Router lets you benefit immediately without rewriting your application. We cut our LLM costs by 30%, while also making our features smarter and faster. We built it for Ramp. Now we’re opening it up to everyone. get access here:
Show more
routers are cool i guess
Happy Wednesday. On today's show: - @travisk (Atoms) - @jasonfried (37signals) - @lqiao (Fireworks) - @maxhodak_ (Science Corp.) - @vral (Ramp) See you on the stream.
Today we’re launching Ramp Router. 3 years ago, we built an internal LLM router at @tryramp that powers AI products for 70,000 customers. Back then it was mostly about saving money. Now it feels obvious: the best model changes constantly. GPT, Claude, Gemini, Grok, Qwen, DeepSeek, Kimi, GLM - prices and capabilities move every week. So we’re opening up access to everyone. One OpenAI-compatible endpoint. The right model for every request. Lower cost without rewriting your app. Reserve access to use it.
Show more
0
216
1.8K
208
Forward to community
The spice must flow
AI is extremely good at spending your money very quietly. our own token spend went from a rounding error to more than 10% of payroll in a year. one week in May we burned through $1.5 m. our CFO didn't love telling me that number, and he really didn't love telling the internet. but every finance leader we talk to is living the same story. the bill keeps going up and teams can't answer basic questions about it. which team is driving it? which models they're using? what changed this month? whether a cheaper model would do the same job? so finance gets two bad options: keep paying and hope, or cut broadly and slow down the work that's actually compounding. we built a third one. Ramp now connects to OpenAI, Anthropic, Gemini, Cursor and pulls it all into one place. see it, understand it, control it - down to a single API key. built with 1,000+ companies managing 100T+ tokens a month, and now spending less than they expected too! try it today at
Show more
starring @justsiddaay @arakharazian @JayaGup10 the token maxxing influencers
🛑STOP SCROLLING🛑 Please enjoy this 2:10 minute history of compute leading up to the present tokenmaxxing era