Guide

Why Your Live Chat Numbers Do Not Add Up

6 minute read · Updated August 15, 2026

Argue about the data before you argue about the target

A recurring meeting: somebody presents a chat number, somebody else disputes it, and forty minutes disappear into whether the target is right. Almost always the target was never the problem. The number was measuring something slightly different from what everybody in the room assumed, and nobody had checked.

Chat data is unusually easy to get subtly wrong, because a conversation has fuzzy edges. It has no clean start, several possible ends, more than one participant, and a duration governed by human behaviour rather than by a system event. Every one of those is a place where a report can quietly diverge from reality.

What follows is the short list of ways it happens, and an audit that takes an afternoon and settles most of the arguments permanently.

Phantom open chats inflate everything downstream

This is the single largest distortion in most chat reporting. A significant share of conversations are never actually closed: the customer stops replying, the window sits there, and eventually something times out.

The consequences depend on how your reporting defines duration, but they are never small. Handle time is inflated by conversations that ended in reality an hour before they ended in the data. Concurrency looks higher than it was, because chats nobody was working on counted as open. Agent comparisons quietly become a measure of who remembers to close things.

The fix is procedural rather than analytical: agents close conversations deliberately, and you find out what your timeout actually is instead of assuming. Any duration metric computed before that habit exists should be read as an upper bound, not a measurement.

Your own team is in the numbers

Test chats are real chats. So are the ones from the developer checking the widget after a deploy, the ones from an agent training a new hire, the ones you started yourself to see whether notifications fire, and the ones from your marketing team looking at the site.

On a high-volume site this rounds to nothing. On a site doing thirty conversations a week it is a material share of the total, and it is biased in a consistent direction: internal chats are answered instantly by somebody who knew they were coming, so they pull first response time down and satisfaction up.

Decide how you will exclude them before you need to. Tagging them is the usual answer, and it only works if it is habitual. The alternative is to accept that low-volume reporting carries a few points of noise, which is perfectly fine as long as it is stated rather than discovered later.

Tags only mean something if they are applied the same way

Tagging is where chat reporting most often turns into fiction. Not because people fail to tag, but because two agents use one tag to mean different things, or two tags for the same thing.

The warning signs are recognisable. A tag list longer than anybody can hold in their head. Near-duplicates differing by a plural or a synonym. A tag applied enthusiastically by one agent for six months and by nobody else. A category that silently changed meaning when a product launched.

The number this produces looks precise and is not. If you intend to make decisions from tag distributions, the list needs to be short, defined in writing and pruned on a schedule. Chat tagging and categorization covers building one that survives contact with a busy queue.

Time zones quietly ruin the time-of-day report

Any report shaped by hour of day is only as good as its assumptions about time zones, and this is the error people find last, because the chart still looks entirely plausible when it is wrong.

Establish three things explicitly rather than by inference. What time zone the underlying timestamps are recorded in. What time zone the report renders them in. And whether both follow daylight-saving changes the way your team's clocks do. An offset that appears for part of the year is easy to mistake for a real change in visitor behaviour, and it will move your staffing decisions if you let it.

The cheap check: pick a conversation somebody remembers having and confirm the report agrees about when it happened. Do it once in summer and once in winter.

One cause of that summer-and-winter disagreement sits in your own account settings rather than in the report. An account that stores a plain UTC offset cannot express whether your area observes daylight saving, so its timestamps can run an hour out for months at a time until a region is recorded — why your timestamps need a region, not just an offset covers what to set and who is affected.

Bot conversations and human conversations are not the same unit

If an AI assistant answers some of your conversations, a total conversation count is now an average across two very different things, and most ratios built on it become hard to interpret.

Response time is the clearest example. An assistant replies in about a second, so a blended first-response-time figure mostly measures what share of traffic the assistant took, not how quickly your people answer. Satisfaction blends differently again, because the two populations are not asked at the same moment about the same kind of problem.

Report them separately, always. And be careful what a resolution or containment figure means: a rate computed over conversations the assistant touched tells you how often a person did not have to step in, which is a containment measure rather than a satisfaction measure. Measuring whether AI chat is working covers what that number can and cannot support.

One person, several conversations

The same visitor comes back three times in a week about the same problem. Depending on the question you are asking, that is either three conversations or one unresolved issue, and a report that counts conversations will tell you the first while the customer is living the second.

This matters most for anything framed as resolution. A high resolution rate sitting alongside high repeat contact is not success; it is the same problem being closed repeatedly. Counting distinct people alongside conversations is usually enough to expose it, and the gap between those two numbers is a more honest quality signal than either one alone.

The afternoon audit

Rather than fixing all of this at once, run a single pass and write down what you find. It is short.

  • Take one week of conversations and count them by hand from the raw list. Compare with what the report says for the same week. Any gap is a definition difference worth naming.
  • Sort that week by duration and read the longest five. If they are phantoms, you have found your handle-time problem.
  • Count the internal and test conversations in the week and express them as a share of the total.
  • List every tag used and mark the near-duplicates.
  • Check one timestamp against a conversation somebody remembers.
  • Split the week into assistant-handled and human-handled, and recompute response time for each.

Write the answers wherever your metric definitions live, so the next person to query a number gets the caveats along with it.

How MyLiveChat fits

The raw material for all of the above is the transcript archive, which is where a hand count comes from and where you settle any argument about what a summary figure is doing. The analytics view is the summary layer, and the point of the audit is to learn how its definitions line up with yours before you build a target on top of them.

If you are pulling chat data into your own warehouse or reporting tool, do the audit first. Every definition problem described here is far cheaper to fix before it has been joined to three other tables. Getting live chat data into your own reporting covers that side of it.

What to measure

  • The gap between your hand count and the report. Ideally zero. If it is not, the difference is a definition nobody has written down yet.
  • Share of conversations closed by an agent rather than by timeout. This is the health check for every duration metric you own.
  • Internal and test conversations as a share of the total. Small and known beats invisible.
  • Number of tags in active use. Rising steadily is the warning sign: it means the vocabulary is being invented rather than applied.
  • Distinct people versus conversations, weekly. The ratio between them is your repeat-contact signal, and it tends to move before satisfaction does.

Ratings deserve their own version of this scepticism, because the survey is withheld from some conversations entirely rather than merely skipped. Why your post-chat survey never appeared covers which conversations are never in the sample.

One further reason two reports disagree is that they are not reading the same table. Where your dashboard numbers are actually computed covers the rolled-up table behind the analytics screen, and the one figure on it that counts AI-touched conversations rather than all of them.

Put it into practice

MyLiveChat keeps every conversation in a searchable archive, so you can hand-count a week and find out what your summary numbers are really doing.

Free forever for 1 agent

Give every visitor an instant way to reach you.

Launch live chat, connect your knowledge base, and add AI answers when you are ready. No credit card, no trial clock.