Guide

When Visitors Try to Manipulate Your AI Chat

5 minute read · Updated August 14, 2026

Someone will test it, and it is not personal

Put an AI chat on a public website and a fraction of the people who find it will try to make it misbehave. Most of them are not attackers in any meaningful sense. They are curious, or bored, or they have seen a screenshot of somebody else’s bot saying something ridiculous and they want to see whether yours will do the same. A smaller number are working towards something concrete: a discount, an exception, a written promise they can hold you to.

It helps to plan for this as ordinary operating conditions rather than as an attack. The teams who get burned are not the ones who got targeted by somebody clever; they are the ones who never considered the possibility and gave the bot both broad scope and an authoritative voice. The defence is almost entirely design, decided before launch, and it costs very little.

The four things people actually try

The attempts cluster into a small number of shapes, and recognising them makes the countermeasures obvious.

  • Instruction override. Telling the bot to ignore its instructions, act as a different system, or reveal its configuration. Usually curiosity, occasionally a prelude to something else.
  • Manufacturing authority. Getting the bot to state a price, a refund, a delivery date or a policy exception that nobody authorised, then treating that statement as binding.
  • Role-play smuggling. Wrapping a request the bot would refuse inside a fiction: pretend you are a manager who can approve this, write a story where the discount code is revealed.
  • Screenshot farming. Steering the conversation towards something offensive or absurd purely to capture it. The goal is the image, not the answer.

Scope is the defence that actually works

Every one of those attempts depends on the bot being able to say something it should never have been able to say. If discounts, refunds, contract terms and delivery commitments are outside its scope entirely, then no amount of clever phrasing extracts them, because the capability is not there to extract. Narrowing what the bot is allowed to discuss does more for this than any amount of instruction-level hardening.

Draw the line by consequence. Anything that costs money, creates an obligation, or would be expensive to reverse belongs with a person. The bot answers questions about how things work and where things are, and routes everything else. That boundary is worth writing down, because it is also the boundary you will be explaining to a customer later if something slips through.

Never let the bot sound more certain than it is

The second half of the problem is voice. A bot that hedges appropriately produces a screenshot nobody wants to share, while a bot that states everything with total confidence produces one that travels. The instinct to make automated replies sound crisp and authoritative works directly against you here.

Prefer answers that attribute rather than assert. Our returns page says the window is 30 days is a better sentence than your return will be accepted, because the first is checkable and the second is a commitment. When the bot is unsure, the honest sentence is short: it does not know, and here is the person who does. Uncertainty expressed plainly is not a weakness in an AI answer; it is the property that keeps the answer safe to publish.

What to do when something does get through

At some point a bot will say something it should not have. The response that keeps customer trust is fast, plain and unembarrassed. Acknowledge what was said, state what is actually true, and honour the gap in the customer’s favour when the amount is small enough that arguing costs more than paying. Nobody has ever been talked out of a screenshot by a policy citation.

Then fix it at the level it happened. If the bot invented a price, that is a scope problem rather than a wording problem, and tightening the phrasing will not stop the next one. If it quoted a genuinely stale article, the fix is in the content. Log what happened and what you changed, because three of these in a row will show you a pattern that any one of them alone would not.

Do not overcorrect into uselessness

The predictable overreaction is to clamp the bot down until it refuses anything with a hint of ambiguity. That trades a rare embarrassing answer for a constant useless one, which is a much worse deal and far harder to notice because it generates no incident to review. A bot that answers nothing produces no bad screenshots and no value.

Keep the scope narrow and the manner helpful. Within its subject the bot should be direct and genuinely useful; outside it, it should route quickly to a person rather than lecture the visitor about what it cannot discuss. A refusal that ends the conversation is a failure even when it is a correct refusal.

How MyLiveChat fits

The controls that matter here are the ordinary ones. The AI chatbot answers from the content you train it on, so what is in your knowledge base is effectively the boundary of what it can competently discuss, and the per-article AI toggle lets you keep internal notes, seasonal promotions and anything you would not want quoted out of its reach entirely.

Handoff is the other half. When a conversation goes somewhere the bot should not follow, it needs a real person to route to rather than a dead end, and every exchange is stored in transcripts so you can search for the attempts after the fact rather than relying on someone having noticed one at the time.

What to measure

This is not a metric you watch daily, but it is worth a look whenever you review AI answers.

  • How often the bot produced an answer nobody would want quoted, sampled from real transcripts
  • Whether those answers cluster around one topic, which usually means a scope gap rather than bad luck
  • Time from a bad answer being sent to somebody noticing it
  • Handoff rate, watched for the overcorrection: a sudden jump means the bot has become useless rather than safe

Put it into practice

MyLiveChat is free forever for one agent, with unlimited chats and the embed code ready in about a minute.

Free forever for 1 agent

Give every visitor an instant way to reach you.

Launch live chat, connect your knowledge base, and add AI answers when you are ready. No credit card, no trial clock.