Register and share your invite link to earn from video plays and referrals.

rahul
@rahulgs
1.3K Following    17.2K Followers
How do you turn an agentic commerce demo into a production-ready agent? With Ramp, you can give an agent its own identity and spending limits, plus the ability to attach receipts, memos, and everything else finance needs to close the books. Here’s what it looks like. We’re onboarding alpha customers. Reach out via the website below for access.
Show more
@vral comment router if you want if you want $26
We tested NVIDIA NeMo Switchyard’s stage router for coding agents in Ramp SWE-Bench. Routed agents showed comparable performance to single-model controls while substantially reducing costs and runtime.
Show more
Contrarians tend to do very well at Ramp. @shevchenkoaalex, who runs @ramplabs, is one of them. High-agency people doing hard work should have real say in what work is worth chasing. Especially when everyone else says "that sounds too hard" or "that'll never work." Those are usually the projects worth trying. If you hear "impossible" and enjoy proving people wrong through action, you're exactly who we want building with us at @tryramp:
Show more
There is no substitute for looking at the data. Manually inspecting failure modes is how you discover the shape of the problem. Great AI products are built by teams who set up their infra + instrumentation to make looking at data easy. @austospumanto @bcherny @rahulgs
Show more
We’re open-sourcing PorTAL, our framework for shared task representations and cross model LoRA adaptation. It now spans from hybrid attention models to multimodal systems including Gemma 4 E2B, Mistral 7B & @thinkymachines' Inkling. Code: ramp-public/portallib Models: @huggingface /RampPublic
Show more
Grok 4.5 is #1# at processing real-world invoices At Ramp, we tested models on 150k bills submitted by actual businesses, scoring them on whether they predicted every correction a human would make Grok achieved the highest perfect-extraction rate, beating similarly priced models Gemini Flash 3.6, GPT 5.6 Terra, and Sonnet 5. This is a demanding long-context reasoning task. The model must infer patterns across 100K+ tokens of prior invoices, business memories, and human corrections, then apply them to new bills. The goal: zero-click accounts payable, with invoices processed correctly without human intervention.
Show more
0
118
1.2K
97
Forward to community
Introducing Genius AI: the platform handling scheduling, payments & admin for service businesses. GlossGenius is now part of Genius AI. We’re nearing $200M ARR, powering over 125K businesses. We’re also announcing our Series D at a $1.15B valuation, with $125M raised to date.
Show more
Every week, the price-intelligence-latency frontier shifts, and we expect this trend to continue Across 100+ use cases in our product, keeping all up to date with the right model is a challenge - Either we're losing out on intelligence for the dollars we spend, or we're spending too much money for the intelligence the feature needs Ramp Router lets you benefit immediately without rewriting your application. We cut our LLM costs by 30%, while also making our features smarter and faster. We built it for Ramp. Now we’re opening it up to everyone. get access here:
Show more
Introducing ask-web: Rox’s in-house web search agent. ask-web sits on the cost-per-accuracy pareto frontier of the hyper-parameter grid when compared to frontier labs and commercial search agent providers. The agent delivers 91.3% accuracy at 1.03 cents per query on real production prompts. It has been running in production for more than 6 months with continuous evals. Inference partners: @togethercompute, @baseten, @modal Commercial Search vendors benchmarked: @perplexity_ai, @ExaAILabs, @p0. Frontier Search vendors benchmarked: @OpenAI, @AnthropicAI Exa, OpenAI and Anthropic excel on accuracy. Parallel and Perplexity are cost-efficient. Here’s the breakdown:
Show more
From spoken spec to working service. Sam Kronick, an engineer on @tryramp’s Applied AI team, used GPT-5.6 to build an entire service end-to-end from a natural-language spec.
now there's nothing stopping you from trying: grok make me a $1 billion dollar b2b saas make one mistake
Every company used to start with paperwork. The next generation will start with a prompt. Ramp for Agents lets AI agents incorporate your company, apply for Ramp, and get your business ready to spend, pay bills, and manage money.
Show more
it’s incredible what swyx has built with aie in just three years
In many ways, finetuning or RLing a custom model is a bet against model progress and scaling. It's to choose to say "we don't think there's going to be a good enough base model for this task anytime soon, so we're not going to wait" with oss release velocity these days, its a hard tradeoff It's easy to end up on a custom model with an outdated base (Kimi 2.6 is only a few months old) So we fixed it - PorTAL lets you swap base models quickly, allowing your learned task specific behaviors to port to new models as they come, no matter how fast
Show more
0
30
1.1K
53
Forward to community
We can finally say AI isn't killing jobs. A new paper from me, @tryramp, and @RevelioLabs uses firm-level spend and workforce data across 21K U.S. businesses to measure AI's impact on jobs. Firms that adopt AI heavily grow headcount 10% over two years following adoption. Low adopters see no statistically significant change.
Show more
0
166
2.7K
613
Forward to community
Today @karimatiyeh and I are both taking new titles as Co-CEOs of @tryramp. If you know us, this won't feel like a change. From when we first started building together twelve years ago, our partnership has run on a couple of motivating principles. On decision-making, we trust each other completely to make critical calls for the company across every function. And on organization design, technology is not a distinct part of the company - it is the entirety of it. That is why Karim has for years directly managed risk, operations, and marketing. Most importantly, at Ramp there is no line between the people who build and the people who do everything else. Everyone is a builder. For the last 2,656 days, we have run the company this way. This only makes it formal. We thought it was important to do it now because of how we see the AI exponential reshaping what Ramp can be. Decisions of company strategy are increasingly decisions of technology and systems design. We have always believed every function should be approached as a systems-engineering problem (even when the system was primarily human) but the rise of machine intelligence makes this existential. Every part of the company must be positioned to leverage the continued explosion in model intelligence and capabilities. If we do this well, each step-change in what models can do compounds automatically into better products and faster execution without anyone having to rebuild the company to capture it. If we fail to operate this way we will ultimately be outcompeted by a new company that does. We are also making Rahul Sengottuvelu our CTO. @rahulgs has led Applied AI at Ramp since joining us three years ago through the acquisition of his prior company, Cohere. Before that, his first company was building customer-service agents on GPT-3 at a time when almost no one knew what a large language model was, and he has spent every year since pushing the frontier of what existing models can do. He has also been right on nearly every major technical direction in AI well before it was obvious. Building Ramp now means applying AI to every part of it, and Rahul is the person stepping up to lead that work. We are still very early in the history of Ramp. Our current chapter is perhaps the most dynamic, but we have never been more optimistic on where it is going and the mission has never been more important. The businesses that trust us are navigating the same shift we are, and we intend to be there for all of it: managing their token spend, supercharging their finance teams, and helping them get more out of every dollar and hour. - Eric & Karim
Show more
0
75
1.1K
55
Forward to community
@tryramp Applied AI is hiring a frontend-leaning engineer for the Financial Intelligence team. Build the AI artifacts 70K+ businesses rely on: agentic analysis, FP&A, month-end close, charts, tables, spreadsheets, and workflows over messy financial data. NYC or SF, details below!
Show more
1. as a mental model it is more correct to think of fable+ class models as english -> code interpreters - converts your idea into code into "correct" code regardless of problem complexity and output complexity (diff size). Fable 5 will be the worst of this new class of models 2. diff size/complexity is to be managed purely for review: small diffs - in high risk areas of code (auth/identity/data access/network access/money movement) large diffs for code that can be empirically verified (frontend/backend plumbing/code without network or db access/performance code that can be empirically verified) 3. time it takes to ship software is completely disconnected from time to produce the PR - how long the work takes depends fully on ability to review/merge code while managing risk at scale 4. solving the bottlenecks for above matter enormously- linters/testing/CI/shadow mode verification/empirical verification 5. agency matters enormously- what are the biggest bottlenecks to speeding up the loop and eliminating them? what are the problems that need solving and when do they need solving? what does it take to the solution to all of them today? 6. deep understanding of the full stack matters enormously- what problems are worth pursuing? is there a higher level of problem abstraction to address first? should I give it the sub-sub task, the sub task, or the task itself. what are the major risks with this PR (order of importance: security holes/correctness holes/performance holes). is there a higher speed way of producing data that allows me to merge this? should this be run in shadow or in a sandbox or a flag. understanding every line of logic may not be needed but understanding and managing risk matters enormously. 7. the cost of complexity itself is changing. it might be now worth "maintaining" 50% more code to get a 5% performance win. getting the right abstractions matter less because larger refactors are less tedious. code quality nits become huge drag. very likely, a much smarter model will be maintaining your code so worth taking on more technical debt now. taking the time to hand architect and rebuild systems comes with an enormous cost of velocity 8. if it quacks like a duck and walks like a duck, it's a duck. For low risk cases, it might be more sane to treat code chunks (services / functions) as a black box, like we do for neural networks: do full empirical verification only: has code produced correct outputs for the last 10,100,1000,10k inputs ? can we quarantine this large piece of code - no outbound access to network / database ? what happens when this code is wrong? do we get hacked/or crash(memory/cpu)/is an inconvenience? is it internal facing or external? what can we do to address these risks? 9. eventually, logical verification (line by line review) will come at an enormous cost- save it for where it matters and build systems that are tolerant to empirical verification. is there a decorator that prevents db / network access? correctness bugs are significantly easier to rectify than access bugs 10. what are the rails that allow for even faster iteration? code permissions can be opt in - db writes, db reads, network egress (to where?), PII access. how long does it take to get shadow mode data? how many PRs can be tested? What are the categories of diffs
Show more
0
65
1.8K
150
Forward to community
yes things are changing fast, but also I see companies (even faang) way behind the frontier for no reason. you are guaranteed to lose if you fall behind. the no unforced-errors ai leader playbook: For your team: - use coding agents. give all engineers their pick of harnesses, models, background agents: Claude code, Cursor, Devin, with closed/open models. Hearing Meta engineers are forced to use Llama 4. Opus 4.5 is the baseline now. - give your agents tools to ALL dev tooling: Linear, GitHub, Datadog, Sentry, any Internal tooling. If agents are being held back because of lack of context that’s your fault. - invest in your codebase specific agent docs. stop saying “doesn’t do X well”. If that’s an issue, try better prompting, linting, and code rules. Tell it how you want things. Every manual edit you make is an opportunity for improvement - invest in robust background agent infra - get a full development stack working on VM/sandboxes. yes it’s hard to set up but it will be worth it, your engineers can run multiple in parallel. Code review will be the bottleneck soon. - figure out security issues. stop being risk averse and do what is needed to unblock access to tools. in your product: - always use the latest generation models in your features (move things off of last gen models asap, unless robust evals indicate otherwise). Requires changes every 1-2 weeks - eg: GitHub copilot mobile still offers code review with gpt 4.1 and Sonnet 3.5 @jaredpalmer. You are leaving money on the table by being on Sonnet 4, or gpt 4o - Use embedding semantic search instead of fuzzy search. Any general embedding model will do better than Levenshtein / fuzzy heuristics. - leave no form unfilled. use structured outputs and whatever context you have on the user to do a best-effort pre-fill - allow unstructured inputs on all product surfaces - must accept freeform text and documents. Forms are dead. - custom finetuning is dead. Stop wasting time on it. Frontier is moving too fast to invest 8 weeks into finetuning. Costs are dropping too quickly for price to matter. Better prompting will take you very far and this will only become more true as instruction following improves - build evals to make quick model-upgrade decisions. they don’t need to be perfect but at least need to allow you to compare models relative to each other. most decisions become clear on a Pareto cost vs benchmark perf plot - encourage all engineers to build with ai: build primitives to call models from all code bases / models: structured output, semantic similarity endpoints, sandbox code execution. etc What else am I missing?
Show more
0
164
5.2K
413
Forward to community