Guide

Testing Changes to Your Live Chat Setup

5 minute read · Updated August 14, 2026

Most chat changes are never actually evaluated

Someone changes the greeting. The widget moves to the left. A proactive invitation starts firing after fifteen seconds instead of thirty. A week later the team agrees it “feels better” and moves on. Nobody can say whether it helped, and if chat volume drifts down three months later, nobody can say which of the nine changes did it.

This is not carelessness. Chat sits at the end of a noisy chain: traffic mix, season, a campaign someone launched, one agent being on holiday. Any of those moves your numbers more than a greeting rewrite will. So the honest position is not that testing is impossible, but that casual before-and-after comparison is close to worthless, and a small amount of structure fixes it.

Decide what you are testing before you change anything

Write down, in one sentence, what you expect to happen and which number should move. “A shorter pre-chat form will raise the share of visitors who finish it” is a testable statement. “This will improve the experience” is not, because no result could ever contradict it.

Name the number before you look at it. If you decide afterwards which metric counts, you will always find one that moved in your favour - there are enough of them. Picking the metric first is the single cheapest discipline in this whole article.

Also write down what you would accept as a failure. A change that raises chats started but lowers the share that reach a useful answer is not a win, and you should decide that before you are emotionally invested in the redesign.

What is actually worth testing

Not everything deserves an experiment. Test the things that are cheap to change, plausibly significant, and easy to reverse:

  • The greeting and the opening question. Wording changes how many people reply and what they say first.
  • Widget placement and prominence. Where the launcher sits, and whether it is visible on mobile without covering anything.
  • Pre-chat form fields. Every field you remove raises completion and lowers the context you start with. That trade is worth measuring rather than guessing.
  • Proactive invitation timing and targeting. The most over-tuned setting on most sites, and the one most likely to annoy people if you get it wrong.
  • Stated hours and the offline message. Telling visitors when you will reply changes whether they bother leaving a message.

Most of these live in your widget customization settings, which means you can change them without touching your website code, and change them back just as fast.

Give the test enough traffic, and enough calendar

Two things ruin chat tests: too few conversations, and too short a window.

Chat volumes are usually small compared with page views, so a difference that looks dramatic across forty chats is often noise. If your change moves a rate from 20% to 25% across fifty conversations, that is four extra chats. It is not a finding.

The calendar matters just as much. Traffic on a Tuesday does not behave like traffic on a Saturday, and the last week of a month does not behave like the first. Run any test for whole weeks - two if you can - so each variant sees the same mix of days. Never compare one week against a holiday week and call the difference a result.

If your site genuinely does not get enough chats to clear that bar, accept it. Run the change for a month, watch for anything obviously bad, and rely on reading transcripts instead of arithmetic. That is a legitimate method at low volume, and pretending you have statistical evidence you do not have is worse than admitting it.

Change one thing, and keep a record

If you move the widget, shorten the form and rewrite the greeting in the same week, you have one result and three possible causes. Change one thing at a time, or accept that you are testing a bundle and can only keep or discard the whole bundle.

Keep a dated log of what changed. It takes seconds and it is the only thing that will let you explain a slow drift six months from now. A line per change - date, what changed, what you expected - is enough. Most teams that cannot explain their numbers simply never wrote down what they did.

Some things should not be tested on visitors

Experimentation has limits that are not statistical.

Do not run a test that leaves some visitors without a way to reach you, or that shows people a longer wait to see whether they tolerate it. Do not experiment with what you say about privacy, consent, or what happens to a transcript - those are commitments, not variables. And do not quietly test whether an AI answer passes for a human; say when AI is answering, every time.

A useful check: if you would be uncomfortable explaining the test to the visitor who landed in the worse variant, do not run it.

Reading the result without fooling yourself

Three habits do most of the work here.

First, look at the number you named at the start, and look at it before you look at anything else. Second, check the guardrail metric - the thing you did not want to break. A greeting that doubles chats started and halves the share that end in a resolved answer has probably just attracted people you cannot help. Third, read a sample of the actual conversations from each variant. Twenty transcripts will usually tell you why a number moved, which the number itself never does.

Be willing to conclude nothing happened. Most changes do nothing measurable, and “no detectable difference” is a real, useful result: it means you can pick whichever version is simpler to maintain and stop arguing about it.

How MyLiveChat fits

Greeting text, widget position, pre-chat fields, hours and proactive triggers are all dashboard settings, so a test is a settings change rather than a deployment - which also means reverting is instant if a variant goes badly.

For the measurement side, chat analytics gives you the volume and outcome counts, and transcripts are where you go to read what actually happened in each variant. Reading the conversations is the part teams skip and the part that explains the result.

One honest limitation worth stating: MyLiveChat does not split your traffic into randomised variants for you. What you can do cleanly is run a change for a defined period, compare like-for-like periods, and keep the log that makes the comparison meaningful.

What to measure

For any chat test, watch a small fixed set rather than everything available: chats started per thousand visitors, the share of started chats that reach a real answer, first response time, and one satisfaction signal if you collect one.

Add the guardrail you chose at the start, and record the sample size next to every number. A rate without a denominator has fooled more teams than any other single thing on a chat dashboard.

Put it into practice

MyLiveChat is free forever for one agent, with unlimited chats and the embed code ready in about a minute.

Free forever for 1 agent

Give every visitor an instant way to reach you.

Launch live chat, connect your knowledge base, and add AI answers when you are ready. No credit card, no trial clock.