Why a published average rarely applies to you
Benchmarks are appealing because they answer a question everyone has: are we doing well? The
difficulty is that almost every published figure is built on a population that does not resemble
your situation, using definitions that are not stated.
Start with the definitional problem, which is larger than most people expect. First response
time can mean the gap before an automated acknowledgement, the gap before a human types anything,
or the gap before the visitor receives a real answer. Those three numbers can differ by an order
of magnitude for the same conversation. A benchmark that does not say which it measured is not
comparable to yours, and most do not say.
Resolution rate has the same problem in a worse form, because it depends entirely on who
decides a conversation was resolved and when. Satisfaction scores depend on when the survey was
shown, how it was worded, and what proportion of people answered — a score collected from
eight per cent of conversations is measuring something quite different from one collected from
sixty.
Then there is the population. Averages across an entire industry blend a company with two
agents and a company with two hundred, a queue of order-status questions and a queue of technical
faults. The average of those is a number that describes nobody in particular.
None of this means benchmarks are useless. It means they are useful as a rough sanity check for
whether you are in a plausible range, and not as a target. Managing towards a published figure is
where the harm starts.
What a number hides
Aggregate metrics conceal by design, and the concealment is usually where the actual problem
lives.
Averages hide distribution. A mean first-response time of ninety seconds is compatible with
nearly every visitor being answered in a minute, and it is equally compatible with most being
answered instantly while a tenth wait ten minutes. Those are different businesses with the same
number, and only the second one has an urgent problem. Looking at the slowest tenth tells you far
more than the mean.
Totals hide mix. Overall satisfaction can hold steady while satisfaction on your highest-value
conversations falls, because volume comes from routine questions that are easy to answer well.
Reading any metric split by topic almost always changes the conclusion.
Period figures hide time-of-day. A weekly average smooths over the fact that your worst hour is
consistently bad, and your worst hour is when a meaningful share of visitors form their
impression. The arrival curve matters more than the weekly total.
And every metric hides the people who never appear in it. Visitors who saw the widget and did
not open it, or opened it and left before typing, are absent from every conversation-based
measure, which means the numbers can improve while the experience gets worse.
Build your own baseline instead
The genuinely useful comparison is against yourself, because the definitions are consistent and
the population is exactly the one you care about.
Write down how you define each metric, in a sentence, before you start collecting. This is
dull, takes ten minutes, and prevents the most common failure, which is discovering six months
later that the definition drifted and the trend is meaningless.
Collect for long enough to see your own variation before drawing conclusions. Most support
queues have a weekly rhythm and a seasonal one, and a fortnight of data will mislead you about
both. Establish what a normal range looks like, then treat movement outside that range as the
signal rather than reacting to every fluctuation.
Segment from the beginning. New versus returning visitors, by page, by topic and by hour are
the splits that repeatedly turn out to matter, and adding them retrospectively is usually
impossible.
Set targets from your own baseline and your own capacity, not from an external figure. A target
of matching a published average is arbitrary; a target of reducing your slowest tenth of waits by
a third is specific, achievable and connected to something a visitor would notice.
When you do reference an external number, treat it as a prompt to investigate rather than as a
verdict. Being well outside a plausible range is worth understanding. Being slightly below an
average of uncertain provenance is not worth a project.
What to measure
Track distribution alongside every average you report — at minimum the median and the
slowest tenth. If you only have room for one number, the slowest tenth is more actionable than the
mean, because it describes the experience of the people most likely to complain or leave.
Keep your metric definitions in writing next to the dashboard, and review them when the numbers
move sharply. A sudden improvement is a change in measurement more often than a change in
performance, and checking that first saves a great deal of misplaced celebration.
Watch response rate on any satisfaction measure, not just the score. A rising score with a
falling response rate usually means the unhappy people have stopped answering, which is the
opposite of what the number appears to say.