Guide

Measuring Whether AI Chat Is Working

5 minute read · Updated August 14, 2026

Start from the decision, not the dashboard

AI chat generates a lot of countable things, and most teams end up reporting whichever of them the tool displays most prominently. Before any of that, settle what decision the measurement supports: whether to widen the AI's scope, whether to narrow it, whether it justifies its cost, or whether it is harming the experience in ways the volume numbers hide.

Those four questions need different evidence, and conflating them is why AI chat reporting so often produces agreement in the meeting and no decision afterwards.

Containment is the headline, and it is easy to fake

Containment — the share of conversations that finished without a human — is the number everyone reports, and on its own it is close to meaningless. A conversation counts as contained whether the visitor got a correct answer and left satisfied or gave up in frustration and went to a competitor. Both look identical in the count, and the second one is a rising number that reads as success.

Never report containment alone. Pair it with something that detects the failure mode: repeat contacts from the same visitor within a day or two, a satisfaction signal on AI conversations specifically, or a periodic read of the conversations that ended without escalation.

Watch the trend rather than the level. A containment rate that climbs while your product and content stayed still is more likely to mean people stopped asking follow-up questions than that the AI got better.

Handover rate is not a failure rate

The second most common mistake is treating every handover to a person as an AI failure and driving the number down. Push it far enough and you get an automation that will not let people reach help, which is the single most reliably hated pattern in customer service.

Handover is the system working when the question is outside scope, when the stakes are high, or when the visitor asked for a person. What matters is the quality of the handover, not its frequency: did the person arrive with the conversation intact, or did the customer have to start again?

Group the handovers by cause and you get a work list rather than a score. Missing content is a content job, contradictory content is a cleanup job, out-of-scope questions are a scope decision, and visitor-requested handovers are usually fine exactly as they are.

Read the conversations nobody escalated

The most valuable AI measurement is not a number at all. Confident wrong answers are, by definition, the ones nobody flagged — the visitor believed them and left — so they are invisible in every metric you have.

Sample them deliberately. A handful of contained conversations read each week, chosen at random rather than by any signal, is the only reliable way to find this class of failure before a customer does. It is tedious and it is the thing that separates teams who know their AI works from teams who assume it.

Keep a running tally of what you find. Two wrong answers in fifty is a very different situation from two in five hundred, and without the denominator a single alarming example either causes an overreaction or gets dismissed as a one-off.

Cost per conversation, honestly counted

Cost comparisons usually pit the model's usage charge against an agent's salary and conclude something dramatic. The honest version includes the work that keeps the AI functioning: content maintenance, weekly review, the handover design, and the time spent investigating when it goes wrong.

Count the human cost of a bad automated answer as well. A confident wrong answer that produces a complaint, a refund or a lost customer is more expensive than the conversation it saved, and a small rate of them can outweigh a large volume of successful containment.

Compare against the realistic alternative rather than a perfect one. The choice is rarely AI or a fully staffed team; it is usually AI or an unanswered chat at nine in the evening, and that comparison is both more favourable and more truthful.

What the MyLiveChat dashboard counts

MyLiveChat surfaces an AI resolution rate over a window you choose, up to ninety days. It is worth knowing precisely what it measures, because the definition determines how you should read it.

It looks at conversations the AI took part in that produced a ticket, and reports the share of those that reached a resolved or closed state with no public agent reply on them. In other words it is a containment measure — the AI handled it and no human had to write to the customer — rather than a satisfaction measure. It counts an agent's internal note as not breaking containment, and it only covers conversations that became tickets.

Read alongside the tool and vendor usage figures reported next to it, that is a fair picture of how much work the AI absorbed and what it cost. It is not, on its own, evidence that the answers were good, which is what the weekly sample of conversations is for.

What to measure

Report four numbers together and resist adding more: containment, handover rate with its causes, a quality signal from actual sampled conversations, and cost per conversation including the maintenance work. Any one of them alone can be moved in the wrong direction while looking like progress.

Set a review rhythm and keep the history. AI performance drifts as your product changes and your content ages, and the useful comparison is nearly always against your own figures three months ago rather than against anyone else's published benchmark.

Agree in advance what result would make you narrow the AI's scope. A measurement programme with no outcome that leads to doing less is not an evaluation, and deciding the threshold before you are attached to the project is much easier than deciding it afterwards.

The handoff record is the other half of this picture, and it is the more diagnostic half. Reading your AI handoff log shows where the assistant stops and why.

Put it into practice

MyLiveChat is free forever for one agent, with unlimited chats and the embed code ready in about a minute.

Free forever for 1 agent

Give every visitor an instant way to reach you.

Launch live chat, connect your knowledge base, and add AI answers when you are ready. No credit card, no trial clock.