Guide

Live Chat Metrics That Matter (And the Ones That Lie)

4 minute read · Updated July 17, 2026

Start from the decision, not the dashboard

A metric is only useful if a bad number changes what you do next week. For a small team running website chat, only a handful of numbers pass that test — and a few popular ones fail it in ways that quietly encourage the wrong behavior.

The four that matter

1. First response time

The visitor's entire experience of "is anyone there?" compresses into this number. Track the median, not the average — one abandoned overnight chat wrecks an average while telling you nothing. If the median creeps up, the fix is almost always staffing or notification configuration, not agent effort.

2. Missed chats

A missed chat is a visitor who asked for help while you appeared available and got silence. It is the single most damaging number on the board because every unit of it is a burned visitor. If it is not near zero, either shrink your displayed online hours to the truth, or let the offline experience (knowledge base, AI assistant, message form) take over honestly.

3. Resolution without escalation

Of the conversations that started in chat, how many ended there — no follow-up ticket, no "we'll email you"? This is the number that says whether chat is a support channel or just an intake form. If you run an AI-first layer, track its share separately: AI-resolved versus human-resolved tells you what the assistant is actually absorbing.

4. Conversation-to-outcome rate

Chats that ended in the thing your site exists for — a signup, an order, a booked consultation. Even measured roughly, it converts chat from a cost center into an attributable channel, and it tells you which pages deserve a proactive invitation.

The two that lie

Total chat volume

Volume rising can mean growth — or a website that got worse at answering questions, a confusing checkout, a broken docs page. Volume is a symptom to investigate, not a KPI to maximize. The goal of a good support operation is often fewer chats about the same traffic.

Satisfaction score alone

Post-chat ratings are answered by a self-selected minority, usually at the emotional extremes. Useful as a smoke alarm; misleading as a performance target — the moment agents are graded on it, the incentive is to avoid hard conversations, not to handle them well. Pair every rating with the transcript before drawing a conclusion.

A weekly ritual that beats all of it

Ten minutes, once a week: read the five worst conversations (longest, angriest, most abandoned). The metrics tell you where to look; the transcripts tell you what to fix — a missing docs page, a canned response that reads cold, an invitation firing too early. Transcript history exists precisely so this ritual is cheap.

Averages hide the experience you are trying to fix

A mean first response time is dominated by the many fast replies and says almost nothing about the visitors who waited. If your average is forty seconds and one chat in ten waits six minutes, the average describes an experience that mostly did not happen to the people who are unhappy.

Report a percentile instead, or alongside. The ninetieth percentile answers the question that actually matters — how bad is it when it is bad — and it moves when you fix the queue rather than when you get more easy chats.

The same applies to handling time and resolution. The interesting conversations are in the tail, and every averaging step you take moves your attention away from them. If you report one number, report the percentile; if you report two, report both and watch the gap.

Segment before you average anything

A single blended figure across every page, hour and conversation type is usually too coarse to act on. Three splits do most of the useful work.

Split by hour of day, because a good daily average can conceal an hour that is consistently poor. Split by entry page, because chats from a pricing page and chats from a support article are different work with different outcomes. Split by whether the conversation reached a human, once AI answers any share of your volume, or your resolution numbers will blend two very different processes.

Resist adding a fourth and fifth split. Beyond about three dimensions the samples get small enough that you are reading noise, and confidently acting on noise is worse than not measuring.

When a metric starts being gamed

Any number attached to individual performance will eventually be optimised directly, usually without anyone deciding to cheat. Response time produces empty acknowledgements. Chats-per-hour produces rushed closes. Satisfaction scores produce agents asking for good ratings.

The defence is not surveillance but pairing. Speed paired with resolution, volume paired with reopen rate, satisfaction paired with a periodic read of actual transcripts. A pair is much harder to move artificially than a single number, because the shortcut that improves one degrades the other.

Watch for sudden improvement with no change in method. A metric that improves sharply while nothing else moves is usually measuring a new behaviour rather than a better one, and the honest response is to ask what changed before celebrating.

One screen deliberately measures nothing, and it is worth knowing why. The live visitor monitor is a snapshot rather than a record, so every figure read off it describes one instant and none of them belong in a trend.

Put it into practice

MyLiveChat is free forever for one agent, with unlimited chats and the embed code ready in about a minute.

Free forever for 1 agent

Give every visitor an instant way to reach you.

Launch live chat, connect your knowledge base, and add AI answers when you are ready. No credit card, no trial clock.