Every AI agent vendor says they're the best.
For weeks, we've been wondering about how to determine who actually is.
Fin used to have a chart on their website showing they are better than Decagon and Sierra. Is that true? It's impossible to validate their claim.
Our goal was to build an objective measurement for our internal teams to know where we're at relative to competitors. Where we're strong and where we need to learn.
So we built a benchmark with AI. 12 agents, 160+ live Shopify and ecom stores, the same typed questions for everyone, a blind LLM grader that never sees the vendor name (we even asked Fable+Astra to tell us if our grader was objective enough to make sure it was!).
At first we thought the ranking would be the interesting part. Who's #
1#, who isn't.
Then we looked at the data and the interesting part was somewhere else:
- "Automation rate" is misleading. An agent can carry a whole conversation, claiming it solved the problem, but actually didn't. Handling a ticket isn't the same as helping the customer.
- The most common failure more is a wall or a loop. Asking for an order number three times, forcing a login, quietly passing everything to a human.
So now, let's talk about the results 🥁
No one wins on all three but Gorgias is #
1# overall
- CX: 🥇
@gorgiasio & Yuma, 🥈 Decagon, 🥉 Sierra
- Sales:🥇 Envive AI, 🥈 Gorgias, 🥉 Fin
- Overall: Gorgias ranks #
1# across all categories
We're running ˜100+ new conversations a day in this benchmark, so results are dynamic.
You can explore the full results here:
Results on our G2 page here:
And if you'd like to do the same thing for your company, we open sourced this benchmark here:
Huge shoutout to
@maxpruv - who spent hours and quite a few tokens getting this live!
Show more
Are you buying an AI agent product? Is your company building one?
If so, how do you answer the following question:
> Who has the best AI agent product in the market? <
We worked on that question internally.
At first, to benchmark ourselves against competition.
Then to understand where are the gaps and what works well.
It took a lot to build:
- Run thousands of conversations
- Building judges that can evaluate a dozen vendor
- Plotting the results and solving for all edge cases
Tomorrow at 11a, we're releasing results and open sourcing our approach to it.
Hope you'll find it helpful to assess agents. Excited to tell you more!
Show more
When you build an AI agent, you need tools to make it powerful.
The thing is, there are 2 types of tools out there.
- The cool ones. They have an MCP, clean API docs, which makes them super easy to plug.
And...
- The not so cool ones. The ones that have no API key, require IT to provision users, or even sometimes a VPN to access. These are really hard to integrate with. You can be creative: reverse engineer the frontend, use a browser agent to perform actions in there. But it breaks, it's tedious, and requires tons of tokens to setup.
As I'm setting up AI agents for brands, my assumption was that hard to connect tools are here to stay, and we'll need to find a way to integrate with them.
BUT
Increasingly, I'm seeing more and more companies ditch backend tool for more agentic era friendly too. Latest example this morning with a merchant migrating to an agentic ERP Fulfil, that offers all the agent friendly tools out there.
I would have never thought that would happen, and I would have expected the ERP category to be the last to be disrupted, but things are moving, and they are moving fast.
Show more
We hit the first 100 customers having Gorgias available on their Slack 🎉
It's amazing to see the merchants asking all sorts of questions.
My favorite ones:
- Asking Cortex (our slack agent) for root cause analysis on ticket trends (why is my CSAT up or down?)
- Asking Cortex to review opportunities to improve AI agent
- Asking Cortex to implement super complex rules
We're opening this beta to a few more accounts.
If you want to have Cortex to manage Gorgias from Slack or Teams, apply here:
Show more
I'm seeing merchants who bought several AI tools in the last few months.
But as we all learn that it's one thing to launch AI agents, and a completely other thing to maintain each of them every simple day, I think consolidation is coming.
Our team has been cooking.
New channels are coming to Gorgias AI agent:
- This summer we added Instagram DMs & Messenger
- Whatsapp is coming in beta next week (hit me up if you're interested)
- And we just added the ability to execute complex actions on SMS
With even more to come later in the year 🙌
Show more
Beta testers wanted!
We launched an agent on Slack that can manage your Gorgias account end to end:
- It QAs tickets daily
- It surfaces opportunities to improve your setup
- It shows you daily stats
We ran an alpha with 20 accounts and merchants loved it. We're launching it to a 100 merchants this week.
Want to give it a try? Apply here and we'll get your Slack account setup:
Lmk if you have any feedback.
Show more
Could an AI agent handle 100% of CX volume for certain brands?
I'm starting to think that's possible.
What I'm seeing:
- The main barrier is trust. Most of the time, AI is missing the mark. But when enough time is invested, enough loops are run, then you can get to a point where the performance is comparable to what a level 1 offshore team would do
- Once that's running, then you need to go down the long list of edge cases. But the good news is, there's a lot of them, but that number is finite. You can work on them one by one, like an Amazon integration, a Target integration, etc.
I think for some brands will soon start having an AI CX manager that manages a CX AI agent responding to customers.
That won't fit all brands (regulated industries, white glove service), but the ones that are looking for a quick way of getting results with "good" quality (but not "great" yet), having an AI run all your CX might be worth it.
If that's something you'd be interested in testing, hit me up, I'd be curious to chat, and try making this happen as an experiment.
Show more
Where do you manage your agents from?
I used to do it from Claude, auditing routine runs first thing in the morning.
Then recently we migrated over to Cortex, our own AI platform and most of my single player routines live here.
But as I'm configuring AI agents for merchants myself, something I'm learning is that these tools are great for single player, but not for multiplayer.
Let's break down multiplayer a bit:
- Multi player is rarely useful by itself
- It's useful when your agent runs the show, and need to ask multiple people for input.
For instance, in the support use cases we see: if a customer's faces an edge case in the warehouse, there's no SOP / common rule you can use. You need to ping the warehouse team, just like a human support agent would do.
That's why I believe increasingly the "harness" or place we'll interact with agents will be Slack (or Teams). Not for everything, but for the majority of interactions, where the agent is asking questions.
I've been testing it with 7 merchants for the past 2 weeks. The pace of iteration is unmatched when the agent can ask questions to the team, and when the team can converse with the agent on Slack.
Because if it's in Claude, it's just not multi player, and it's usually only one way (you talking with the agent).
We're adding more and more Gorgias customers to a Slack channel with us where Cortex can run the implementation / optimization process. Excited to learn from this with more scale!
Show more
Should you build or buy your agent?
It's a hard call to make.
I was chatting with a merchant this week. They run a big business, with very thin margins. They are running on our helpdesk but built their own AI agent.
We discussed it together, using the same latest OpenAI model, with solid QA, well integrated into their system, they built something solid.
I took up the challenge of selling them our own AI agent. At first I thought: these guys are hyper optimized for cost, given how much checks+evals we run on our own, we won't be able to make the ROI worth it.
Then I discovered some interesting points:
- Their agent was escalating to humans more than it should. Even with offshored hired directly at $4 an hour, that cost added up. Cost: ˜15k a month
- Customers were getting upset when the AI hallucinated. Cost: 5k a month in chargeback
- AI is fun to build (I'm addicted to hit with my 100+ sessions a day!) but running it in prod every day creates hidden maintenance work you don't see at first. Cost: usually less than buying here
At the same time, like I shared a few days ago on building our own internal AI platform for Gorgias, when your use case is very specific to your business, build makes a ton of sense.
But what I'm finding for AI agent on ecom customer support on Shopify, there's enough similarities between brands to make the math more favorable on the buying side.
Show more
Big day today for everyone who's using our Helpdesk: we're making Gaia available to ALL CUSTOMERS (including the ones that don't use AI agent).
Gaia is your AI companion in Gorgias, that makes your day 2x more efficient.
Here's how people use Gaia:
- Drafting & editing customer replies
E.g. 'Draft me a warm reply and say that I've refunded the 100 points for the unused $5 voucher"
- Product help & analytics interpretation
E.g. "How would I integrate facebook into gorgias"
- AI Agent setup & skill tuning
E.g. "find gaps in this skill's instructions"
Hope you'll find it helpful, more on this video by Felipe Mora ⤵️
Show more
I've been running an interesting experiment the last 14 days.
I've focused heavily on churn and starting working on re-implementing customers that got stuck on AI agent, getting hands on with customers.
I started with a customer who'd spent months in onboarding, built 14 custom actions, but never went live. The reason? A trust standoff: they wouldn't turn the AI on until it was proven.
So I started working on solutions for them with the help of our team
1. Testing before going live. We built a test suite of 31 scenarios and ran simulations through Cortex (our internal AI agent). Each time a simulation would fail, Cortex would trigger trigger our merchant facing agent Gaia to fix them. Running this loop of autopilot compressed days of work in a few hours.
2. Tone of voice. The customer had an 10+ pages tone of voice policy. Instead of manually tweaking it over weeks, I used a similar approach, simulate hundreds of conversation to find the exact right wording of the tone of voice prompt.
3. The long tail of requests. Custom widgets, integration fixes, edge cases, you name it. This stuff that eats lots implementation time. So I put Cortex directly in the shared Slack channel with the merchant. When they asked for order hold/cancel/resubmit buttons, Cortex read the API docs, confirmed the endpoints, and built the configs.
Taking these learnings home:
- We're prioritizing trust building. You can now run simulations via the MCP (you might need to reload the latest tools), and it's coming to the product in Sept.
- We're accelerating the pace of iteration, putting Cortex on more re-optimization and implementations of customer. The goal is that each implementation makes the next one easier by compounding learnings.
5 of 6 merchants are now live with this setup. We're treating implementation and re-optimization as a product, with a PM who builds, tests, and ships.
When you're behind on a metric, the fastest way to solve it for me is usually to get my hands dirty, using our own product, and accelerating customer feedback loops. Excited to now scale this approach further!
Show more
Should you build an MCP or an AI copilot (aka at Gaia on Gorgias) in your app?
In Q2, we launched both, now that it's been a few months, it's a good time to share some of the learnings here.
- First, usage is high. More than half our customers have used our internal AI tool, and close to half on the MCP.
- The MCP was perceived as a tool for "technical users". A few weeks in, I'm seeing more an more non technical users adopt it. They usually describe it as a game changer, making their job way easier.
- It drives adoption of hard to use features. We've had rules (our workflows) since day 1 on Gorgias, we've tried to simplify them, but very often they required heavy hand holding from our support team.
- Coincidently (that might be unrelated but it's still an interesting observation), our support queue is now close to 0. Our support team has gone more efficient for sure, but I think these AI tools are also reducing support
- The median usage on the MCP is 2x the AI copilot, meaning heavy workflow like bulk exports happen more on the MCP.
- It's still the beginning. This week I shared the MCP with 2 merchants who hadn't heard of it and were blown away
If you haven't used Gaia on Gorgias, click at the bottom left of your account on the "ask Gaia" button (or command G).
For the mcp, you can install it here: If you have several Gorgias accounts, you can connect all of them.
Excited to see CX managers being 10x more efficient with these tools 🔥
Show more
It's been one month we've moved to usage based pricing on Claude, which in my case costs 6x more than the previous subscription cost, so I wanted to share some of the learnings I had through this period.
BEFORE AUGUST (everything on Claude world)
- We had packaged our company skills and MCPs on github with the whole company contributing to these
- I would use max out claude weekly limits each week and use opus as my daily driver
- I would hit limits and had the need to connect user mcps to run certain tasks, for instance truly doing email (vs just creating drafts), doing slack from claude on shared channel, properly doing google slides, etc.. Some of these MCPs could only run on my laptop
NOW
- Our internal AI team has done a wonderful job building an internal platform that is connected to all the skills, MCPs we need. It loads all our skills, all the necessary MCPs
- We are able to tweaks MCPs the way we want and host them in the cloud. It outperforms Claude when we tag it on Notion comments or mention it on slack for my use case
- We are able to plug routines with the exact context we want via slack apps
- Everything is multiplayer and cloud first, which is better for collaboration + work on the go
- We can easily experiment building agents that are customer facing in a multiplayer way too
- Cost remains in control, you choose the model, it's a mix of Kimi 2.7 for basic tasks and frontier models for more complex ones
- Where there's something I want to change, I can PR into it and get the change live within hours
Net net, I'll miss what felt having unlimited opus access, but the ease of use of having your own internal platform really well connected to our own tools is a game changer.
More on this here:
Show more
I worked on 5 implementations last week and got some fresh learnings that might be helpful to you.
The main blocker is building trust.
Will the agent reship the right items?
Will someone find a way to hack it and get refunds?
We run lots of evals on AI agent, but until now we've been missing merchant level evals. You can manually test things in the playground, but it's been too much work to run.
So we've added the ability for the MCP (and soon Gaia) to run these tests in batch and produce results on ˜100 tickets.
This way, each time you make a change, you can run these evals, and get a sense for progress you're making.
Then the magic is that you can iterate on this as a loop, and improve your skills running several loops.
Try it now on the MCP, if you're open to sharing results, I'm keep to learn too!
Show more
If you're an agency managing multiple Gorgias accounts, worth noting that you can now manage them all from your AI app connecting *multiple* Gorgias MCP to the same session
Details on how to do it here: using for each brand you connect.
Core use cases for this include:
- Generating reporting for all the brands you managed
- Implementing a similar playbook across all your brands, for instance updating AI agent skills
- Adding users to multiple accounts all at once
Let us know if you need us to build more tools for it.
Show more
We talk very little about how hard it is to implement AI agents.
Most of the conversation is focused on *why* we should use them, but the *how* is by far the hardest part.
I'm spending my whole week on streamlining how to get to a top level experience and will share learnings soon.
The best example I know of a brand that nailed their AI agent is Mia Chapa at GLAMNETIC.
They achieved amazing results: 55% automation rate (and 65% on good days!) with an astonishing 4.8 CSAT.
Since it's more about the journey than the destination, we sat down in LA to record how working together, we got there, for the 2nd episode of our behind the inbox series.
Check out the story of how they got to what I think is one of the best AI agents in ecom, and hopefully you'll learn a thing or two what it takes to get there and how you can replicate the learnings.
Now the next challenge for us is to lower the bar to get there. The Decagon CEO had a good post on this here earlier this week It's all about productizing the experience.
Gaia helps a lot there, and stay tuned for more on Gorgias helping you get an amazing AI agent with minimal effort!
Show more
In their earnings call last week,
@Shopify shared that AI traffic in ecom is up 3x YoY and converts 2x better.
Amazon now displays Ask Alexa on desktop by default.
As more shoppers want to chat to find the products they want, we made some improvements to our shopping assistant.
It now fetches way more product attributes:
- tags
- categories
- collections
- stock
- variants
- images
When you give more to the AI ... the AI delivers. We're seeing +19% more add to carts on AI convos.
Try it out if you're using shopping assistant on Gorgias, hope you'll like the new recs. If you're not, you can test on:
Show more
I was chatting with a $100M GMV brand yesterday about blockers to automate more with AI.
They faced a roadblock: it's truly hard to integrate their stack with Gorgias' AI agent.
That's a topic we spent time discussing with Chris Long at the last DotDev, and there's a lot that we can do working with partners to solve this issue.
So we're releasing 51 new actions for AI agent:
- Subscriptions: cancel, pause, resume, skip delivery, update billing date and shipping address with Smartrr (on top of existing ReCharge and Skio integrations)
- Returns: initiate, check eligibility, retrieve status and details, add notes, flag for review, and cancel return requests with Loop, ReturnGO, Redo, and Baback
- Loyalty: check points and VIP status, award or remove points, retrieve rewards, and redeem rewards with LoyaltyLion, Rivo, and Okendo
- Store credit: check balances, issue goodwill credit, create gift cards, and deflect refunds with
- Shipping: retrieve shipment information and tracking events with Sendcloud
More to come soon. Thanks Fraser Bruce for making this happen 🙌
Show more
We've heard several requests by Enterprise customers to be able to better track how their CX team spends their time.
So we've implementing more granular statuses, so that you can have visibility into what your team spends time on.
Benefits:
- See where non-ticket time goes
- More gracious wrap up for agents on calls with customers
- Know who's reachable now
- Helps with planning too
Hope you'll find it helpful!
Show more
We’ve raised over $100M to create AI customer support that doesn’t sound like AI slop.
Introducing Gorgias AI Agent 3.0: the first AI customer support that’s better than any human.
How it works 👇
Show more