Register and share your invite link to earn from video plays and referrals.

Augusto Marietti | API father
@sonicaghi
CEO @kong. No AI without APIs. American dream immigrant ๐Ÿ‡บ๐Ÿ‡ธ/๐Ÿ‡ฎ๐Ÿ‡น.
1.1K Following    5.7K Followers
We proxy trillions of API requests every day for the largest organizations but @kong AI token traffic has enter a whole new level of exponential growth. APIs <> Agents <> LLMs. No AI without APIs.
Show more
Private equity is the digital graveyard of technology.
Gartner just released a new Magic Quadrant, and it's forcing the industry to answer a tough question: What does AI governance actually mean? In the new MQ, Gartner formalized AI governance as a distinct enterprise buying category, and they project the market will be worth $1.4 trillion by 2030. But we have to be precise about what this category does and does NOT include. As outlined by Gartner, AI governance platforms are built for CISOs, compliance officers, legal teams, and risk functions. Their job is to manage things like dynamic risk scoring and compliance framework mapping (EU AI Act, NIST AI RMF, ISO 42001). This is the "what" part of AI governance. But it doesn't cover the "how". Gartner is explicit about this point: governance platforms do NOT enforce policy in isolation. They depend on something beneath them to make those decisions operational at runtime. That's the "how" layer, where @Kong lives. Applying AI governance at the traffic layer. Rate limiting, access controls, prompt inspection, PII sanitization, content filtering, etc. This is the enforcement infra that makes governance decisions scalable. It's like traffic law vs traffic lights. You can set broad policies, but you need the traffic layer enforcement to make it actually work. A policy that says "no PII crosses this boundary" does nothing until something in the request path actually checks and enforces it. So what is AI governance? It depends on who you are. CISOs can focus on the "what" layer, while builders need to obsess over the "how". Orgs have to treat governance and AI connectivity as complementary infrastructure decisions. One layer defines the rules. The other makes them real. Traffic law AND traffic lights.
Show more
Token spend becoming a problem? An AI gateway at the traffic layer gives you 3 levers that can drastically reduce token consumption. 1) Prompt Compression: strips unnecessary characters from a prompt before it ever reaches the foundation model. 2) Semantic Caching: caches responses based on meaning, not exact wording, so duplicate intent doesn't trigger a redundant model call. 3) Semantic Routing: routes prompts to lower-cost models based on intent, reserving expensive models for complex tasks and cheaper ones for simple requests. These are valuable at any size org, but at enterprise scale (millions of daily requests) the token cost optimization could reshape your budget entirely.
Show more
MCP all the things!
Announcing the hosted X MCP. Agents now have access to the best real-time information source in the world. Connect Grok, Cursor, or any MCP-compatible AI tool to the X API without any setup! Check it out here:
Show more
Who says commutes arenโ€™t fun?!
First day at our new Austin office. Video of my 5 minute jet ski commute to work.. ๐Ÿ˜๐Ÿ˜…๐Ÿ˜‚
why many of the largest enterprises use @kong AI gateway for token cost management.
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) โ€“ Engineers can choose any model they want, but defaults matter. Weโ€™re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work. Better Routing โ€“ In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task. Better Caching โ€“ Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so weโ€™re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% โ†’ 60% in LibreChat once properly implemented. Keep Context Lean โ€“ Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted. Better Visibility โ€“ Our engineers can use as many tokens as they want, from whatever model they want, but weโ€™ve made usage visible โ€“ and the more you spend on AI, the more impact we expect. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.
Show more
Micron is growing faster than Nvidia. Since a kid I knew that recovering the memory of @RoboCop was the real deal.
MCP: new enough to be exciting, old enough to create a sprawl problem. Tell me if this sounds familiar. One team spins up an MCP server. Then five more. All built differently of course! Nobody knows what's out there or who owns what. Yikes. Agents got flooded with tools they don't need and you accidentally spent a months worth of tokens in a day (and you can't even blame Fable). Shadow infra at it's finest! But now it's got AI speed, so it's like supercharged chaos. That's why we built a central gateway for all AI context. One place to enforce authentication and security for all your data and tools. One point of observability to see which tools are being used, by which agents, and at what cost.
Show more
Increasingly, I believe companies may need to be rebuilt from the ground up, where you have a single timeline of all observability + product metrics + file changes laid out in a retrievable system, like Datadog + Posthog + Google Drive + Slack (really unified filesystem of Claude Code chats + Codex chats). This might be the new data foundation for any and all companies to maximize AI. Needs to be rebuilt because keeping track of diffs on existing system basically impossible to produce longitudinal information on decisions and rollbacks, something coding agent storage companies are actively trying to figure out, but this should extend to businesses as a whole. Highly skeptical existing businesses will adopt this though because it means overhauling everything about their instrumentation and business data, but I think businesses built on this foundation probably can execute 100x better and faster
Show more
0
210
2.3K
180
Forward to community
"The brain is part of intelligence, but there is all our nervous system, peripheral and central. Tool use is what makes intelligence go into a peripheral nervous system." @sonicaghi on the next step change for frontier models: "Things are getting smarter and smarter in so many dimensions, agentic coding, knowledge work, legal, computer use, biology, cybersecurity." "But tool use... the ability for a model to call the right MCP server or the right APIs to get something done, it's pretty limited even with the latest model." "Raw intelligence is improving. The ability to connect the dots is not improving much."
Show more
Every SaaS company will just become an API key.
What happens to SaaS when agents do all the clicking? @sonicaghi, CEO and co-founder of Kong. "Most of the flashy SaaS companies will just become an API key in somebody's .env file embedded into agentic coding loops." "Think about Salesforce... Marc Benioff just announced Atlas 360. A $40-50 billion revenue business rebuilding itself to be agentic consumable. All MCP and APIs. All programmatic access for agents." "What Intercom did with Fin was the playbook: go headless, just an API. Reuse the same data, the same customers, the same key features, but expose them as APIs."
Show more
Fun to be back on @MTSlive today to discuss Cursor acquisition, Fable, and the importance of an AI context layer. Lots of AI news this week!
LLM providers are trying to build enterprise guardrails directly into their products, but it won't work. Companies don't run on a single LLM any more, they're often using a mix of frontier models plus open-source and specialized vertical models spread across various use cases. Model providers can improve their guardrails, but that won't fix the fragmentation problem. The industry tried this before! In the early microservices era, every team baked auth, rate limiting, and logging directly into their apps, which worked fine until you had hundreds of services and no consistent enforcement or unified observability. We abstracted the connectivity logic to the traffic layer back then, and you need to do it again with AI. Effective AI governance (token budgets, permissions, compliance, etc) can't live inside a single model. You need a neutral layer that sits in front of everything. A "Switzerland for AI" as Khozema Shipchandler called it. A single control tower that manages all traffic regardless of source.
Show more
All saas will become an api key
Welcome Salesforce Headless 360: No Browser Required! Our API is the UI. Entire Salesforce & Agentforce & Slack platforms are now exposed as APIs, MCP, & CLI. All AI agents can access data, workflows, and tasks directly in Slack, Voice, or anywhere else with Salesforce Headless 360. Faster builds, agentic everything. ๐Ÿš€ #Salesforce# #Agentforce# #AI#
Show more
That's the America I always wanted to be part of and the reason I moved here. ๐Ÿ‡บ๐Ÿ‡ธ๐Ÿซก
SpaceX created over 4,400 millionaires today, and many of them are regular working people. They are not executives or founders, they are welders, technicians, machinists, and launch crew, the people who showed up every day and built the rockets with their hands. Around 400 of them are sitting on stakes worth over $100 million each. For context, Google's IPO created roughly 1,000 millionaires. Facebook's created around the same. SpaceX is doing more than four times both of them in a single day. Juan Hernandez is one of them.
Show more
People live the present and often forget how long it takes. The @SpaceX grinding has been going on since 2002!!! Founders that keep at it, will compounds over decades into something far bigger than their original dream. If you're going thru hell keep going. Congrats $SPCX crew.
Show more