Measuring chatbot ROI: the metrics that matter
Deflection rate, containment, and revenue per conversation — which numbers are worth reporting, and which ones flatter you.
Chatbot dashboards are designed to make chatbots look good. That is not a conspiracy, it is a product decision: the vendor knows which numbers cause renewals. Containment goes up and to the right, deflection is expressed as a percentage, and a savings figure appears that nobody in finance would accept.
The metrics that tell you whether the bot is actually working are less flattering and considerably more useful. Here is what to track, and — just as importantly — what each number can hide.
Containment, and how it lies
Containment is the share of conversations the bot handled without escalating to a human. It is the headline number on every dashboard, and it can be improved instantly by making it harder to reach a person.
Any metric that improves when you make the customer's life worse is a metric you must never optimise on its own.
Containment is a diagnostic, not a goal. It is only meaningful when read alongside two other numbers: abandonment and reopen rate. Containment up, abandonment up means people are giving up, not being served. Containment up, reopen rate up means the bot is closing conversations rather than resolving them.
Abandonment — the number nobody reports
Abandonment is the share of customers who start a bot conversation and simply stop replying. No escalation, no resolution, no complaint. They just leave.
It is the single most honest signal you have, and it appears on almost no vendor dashboard, because it is the number that makes bots look bad. Instrument it yourself if you have to: any conversation where the bot sent the last message and the customer never returned is a candidate.
Reopen rate
A conversation the bot marked resolved, which the customer reopens within a few days, was not resolved. It was postponed, and it will now cost you a human anyway — plus the goodwill you spent making them ask twice.
Reopen rate is the honesty check on containment. A bot with 60% containment and a 5% reopen rate is doing real work. A bot with 80% containment and a 30% reopen rate is a very efficient way of annoying people.
First response time, split by path
Report first response time twice: once for conversations the bot answered, once for conversations that reached a human. Reporting it as a single average lets an instant bot reply hide a two-hour human queue behind it.
In practice, the blended average is the number most likely to make a team complacent. Every conversation the bot answers instantly drags the mean down, and the customers waiting two hours for the escalation — the ones with the real problems — disappear inside it.
Cost per resolved conversation
This is the number finance will actually accept, and it is the one vendors are least keen to compute, because it requires acknowledging the cost of the bot itself.
- Take the fully loaded cost of the automation: licence, build time, and the ongoing maintenance nobody budgets for.
- Add the platform's conversation charges.
- Divide by conversations genuinely resolved — excluding abandoned ones and excluding those later reopened.
- Compare against the same figure for human-handled conversations.
Most bots come out well on this once they are stable. The point is not to prove the bot is expensive — it is that a number computed this way cannot be gamed by tightening the escape hatch, which is exactly why it is the one worth reporting upward.
Revenue per conversation, for the bots that sell
If the bot qualifies leads or recovers carts, it belongs on the revenue side of the ledger rather than the cost side, and the arithmetic changes completely.
- Attribute revenue to conversations the bot touched, with a defined window — thirty days is a reasonable default.
- Hold out a control group that receives no automation, so you can see what would have converted anyway.
- Track revenue per conversation, not total revenue, so growth in volume does not disguise a fall in quality.
The holdout is what makes this credible. Without it, every bot in history has claimed credit for revenue that was going to arrive regardless, and everybody in the room knows it.
CSAT, measured separately
Ask for a rating at the end of every conversation and report bot-handled and human-handled separately. A blended CSAT is nearly useless, because it lets a strong human team carry a weak bot for months.
The pattern to watch for is a bot CSAT that is acceptable on average but bimodal underneath — lots of fives from people whose question it answered, and lots of ones from people it trapped. The average looks fine. The ones are your churn.
The scorecard
Six numbers, reported together, none of them alone.
- Containment — how much it handled.
- Abandonment — how many it lost.
- Reopen rate — how much it only appeared to solve.
- First response time, split by bot and human.
- Cost per genuinely resolved conversation.
- CSAT, split by bot and human.
Read together, they are almost impossible to game. Read individually — and containment in particular — they will let a bot that is quietly damaging your customer relationships look like the best investment you made all year.
Report all six. If the bot is good, it survives the scrutiny easily. If it does not survive it, you have learned something considerably more valuable than a flattering percentage.
Building the business case, before you buy
Most chatbot business cases are written backwards: somebody has decided to buy the bot, and the spreadsheet is assembled to justify it. The numbers are always enormous, and they are always built on a deflection percentage supplied by the vendor.
A defensible case starts from your own inbox instead, and it takes an afternoon to build.
- Count conversations per month, and sort them by question type.
- Take the top three types and ask honestly which are genuinely automatable — factual, answerable from data you hold, unambiguous.
- Multiply that volume by your true cost per human conversation, including the salary, the tooling, and the management overhead. Not just the salary.
- Discount it. Assume the bot handles perhaps two-thirds of what you think it can, because it will.
- Now subtract the cost of the platform, the build, and — the line everybody forgets — the ongoing maintenance.
The number you are left with is usually still good. It is simply a third of the one in the vendor's deck, and it will survive a conversation with your finance team, which the vendor's number will not.
The soft returns, which are real but hard to bank
Not everything a bot does well shows up in a cost model, and it is worth naming these separately rather than smuggling them into the ROI number as invented pounds.
- Coverage at hours you do not staff. A bot answering at 2am is not deflecting a ticket — it is capturing a customer who would otherwise have gone elsewhere by morning.
- Consistency. Every customer gets the same accurate answer, which is a claim no human team can honestly make on a Friday afternoon.
- Agent retention. Removing the most repetitive third of the job is a genuine improvement to it, and turnover is expensive.
- Data. Every automated conversation is structured by design, which means you finally know what people actually ask.
Report these as benefits, not as numbers. The moment you assign a currency value to 'improved consistency', you have written a business case nobody senior will believe — and you will have undermined the parts of it that were solid.
When the honest answer is to switch it off
Sometimes the numbers come back and the bot is not working. This is a legitimate outcome and it is worth planning for, because organisations are extraordinarily bad at killing things they have already paid for.
The sunk cost of the build is not a reason to keep a bot that is losing you customers. It is simply the price of finding out.
The signals that it is time: abandonment climbing while containment climbs. Reopen rate above roughly a fifth. Bot CSAT materially below human CSAT and not improving after two rounds of fixes. Agents who describe the handoffs as making their job harder rather than easier.
If you see those, do not scale it, do not add features, and do not commission a redesign. Narrow it. Take it back to the single question it handles well, switch off everything else, and let it earn its scope again one job at a time. A bot doing one thing perfectly is a genuine asset. A bot doing eight things adequately is a liability with a very good dashboard.
Presenting this to a board
The six metrics above are the right ones internally. They are not the right ones for a board meeting, where you have four minutes and an audience that will remember exactly one number.
Lead with cost per resolved conversation, compared against the human equivalent, with the maintenance cost included. It is the number that cannot be gamed, and it is the one a finance director will engage with rather than tolerate.
- One slide, three columns: before, after, and the delta. Not a dashboard screenshot.
- Show the CSAT split — bot and human, side by side. Volunteering the uncomfortable number buys you credibility for the flattering one.
- Name the risk honestly: the bot degrades if nobody maintains it, and here is the person who does.
- Do not project savings from deflection you have not yet achieved. Report what happened, not what the vendor says will happen.
The measurement mindset that actually holds
Every metric in this article can be gamed, and most of them will be, usually without anybody intending to. Containment rises because somebody buried the escape hatch. CSAT rises because the survey only reaches people who got an answer. Cost per conversation falls because the definition of 'resolved' quietly loosened.
The defence is not a better metric. It is the habit of reading actual conversations.
Twenty transcripts a month will tell you more about whether your bot is any good than any dashboard ever built. Nobody wants to hear this, because dashboards scale and reading does not.
So do both. Track the six numbers, present the one that survives scrutiny, and read the transcripts anyway — because the moment the automation starts failing, it will show up in the conversations weeks before it shows up in the chart. And by the time the chart moves, a few hundred customers have already decided, quietly and permanently, that dealing with you is more effort than it is worth.
Three more,
worth your time
Ready togrow on WhatsApp?
Join 10,000+ businesses turning conversations into revenue.
Start free trial→Book a demo
