QA is coaching, not surveillance
Quality assurance goes wrong the moment it feels like a hunt for mistakes. Agents who believe QA exists to catch them get defensive, game the scorecard, and stop taking the small risks that make chat feel human. Done right, QA is a shared vocabulary for what good looks like and a steady stream of specific, kind feedback. Frame it that way or it will backfire.
Keep the scorecard small
A rubric with twenty criteria gets filled in mechanically and teaches nothing. Pick the handful of things that actually define a good conversation on your team — did we solve it, was the tone right, was it accurate, did we set clear expectations — and score those. A short scorecard you use every week beats a comprehensive one you abandon in a month.
Sample, don't audit everything
- Review a handful per agent per week, chosen to include a mix — not just the disasters or the wins.
- Deliberately sample low-rated chats, where the lessons are richest, alongside random ones.
- Rotate who reviews, so the standard is shared rather than one person's taste.
Score the conversation, coach the person
The scorecard rates the transcript; the conversation with the agent is where improvement happens. Lead with what went well — specifically, not as a throat-clearing before the criticism — then pick one thing to work on. One concrete, improvable habit per session sticks; a list of ten faults just demoralizes.
Look for system problems, not just agent problems
When the same issue shows up across agents, it is not an agent problem — it is a missing canned response, an unclear help article, a product rough edge, or a policy with no fast path. QA that only ever produces "the agent should have…" is missing half its value. Feed the patterns back into your content and process, not just your coaching.
Make good work visible
Share great transcripts with the team, with names on them and permission asked. Nothing teaches "this is our standard" faster than a real example everyone can see. In MyLiveChat, transcripts are searchable and reviewable, so building a small library of exemplars — and the conversations to learn from — costs attention, not tooling.
Calibrate the reviewers before you trust the scores
The failure that quietly invalidates a QA programme is reviewer drift. Two people score the
same conversation differently, agents notice, and the scores stop meaning anything — at
which point the exercise costs time and produces resentment instead of improvement.
Calibration is the fix and it is not expensive. Have everyone who scores review the same two or
three conversations independently, then compare and discuss the differences. The point is not to
agree on a number but to surface why one reviewer marked something down and another did not,
which almost always reveals that a criterion means different things to different people.
Run it regularly rather than once at setup. Standards drift as the team gets busy, new
reviewers join, and the product changes. A short calibration session each quarter keeps the
scores comparable over time, which is the only way trends mean anything.
Write down what each criterion means with an example of a pass and a fail. “Was the tone
appropriate” is unscoreable as written; the same criterion with two real transcript
excerpts attached is scoreable by anyone. This document is the actual QA programme —
everything else is process around it.
Review the conversations that teach you something
Random sampling is the right default because it is unbiased, but a purely random sample wastes
much of the effort on conversations that were fine and had nothing to teach. A better approach
mixes a random baseline with deliberately chosen cases.
Worth reviewing on purpose: conversations that ended abruptly, conversations where the visitor
came back about the same issue within a few days, conversations that were escalated, and the
longest ones. Each of these has a higher chance of containing something instructive than an
average interaction.
Also review the very good ones, and share them. Most QA programmes only surface problems, which
teaches the team that being reviewed is a threat. Circulating one excellent conversation a month
with a note on what made it work does more for quality than several corrections, and it costs
nothing.
Keep the volume honest. A target of reviewing everything guarantees the reviews become shallow
tick-boxes; a small number reviewed properly, with real written feedback, changes behaviour.
Where time is short, review fewer conversations rather than reviewing more of them badly.
What to measure
Watch the spread of scores, not just the average. A programme where nearly everyone scores
highly is usually measuring nothing — either the criteria are too easy or reviewers are
avoiding conflict — and it will not detect a real decline when one happens.
Track whether feedback changes anything. The honest test is whether an agent who received
specific coaching on a criterion scores differently on it two months later; if not, the feedback
is being delivered but not landing, and the fix is in how it is given rather than in the scoring.
Finally, compare your QA scores against what customers say. When internally excellent
conversations produce unhappy customers, the scorecard is measuring the wrong things, and the
scorecard is what should change.