When can an AI agent reply to customers on its own? A practical UK SME guide to approval rules, data access, testing and a safe route back to a person.

An AI agent can prepare a customer email. Whether it should send that email without a person checking it depends on what the message can change. I would start with drafts, approve every send, then consider automatic replies only for narrow, tested questions where a wrong answer has a limited consequence. Complaints, refunds, prices, promises and sensitive cases should reach a person.

The decision is about the whole customer interaction. A reply that sounds polite can still quote the wrong price, miss a cancellation right or expose information from another account. The send button is a business action, even when a model writes the words.

Why does the business remain responsible?

The Competition and Markets Authority’s guidance on AI agents is clear: consumer law applies whether a customer deals with a person or an AI agent. The business remains responsible for the agent’s actions, including when a third party provides it. The guidance calls for accurate information, testing, monitoring and human oversight.

That does not make every automated message a bad idea. A straightforward acknowledgement that says when a person will reply is different from an agent deciding whether someone deserves a refund. The practical job is to draw that boundary before connecting AI to an inbox.

The National Cyber Security Centre’s interim advice asks operators to define what an agent may do, when it must stop for approval and what technical controls enforce those limits. It specifically warns against relying on a prompt alone. A sentence telling the model to be careful cannot prevent a connected tool from sending the wrong message.

Which customer emails need approval?

I would divide the work into three lanes. They are operating choices for a first deployment, not a legal classification of every email.

LaneExampleWhat the AI may do
Draft and approveA product question, delivery query or routine follow-upFind relevant approved information and prepare a draft. A person checks the recipient, facts and tone, then sends it.
Bounded automatic replyAn acknowledgement of receipt or a status answer pulled from a verified fieldSend only an approved template or tightly limited answer when the required data is present and the case matches the tested rule.
Stop and hand overA complaint, cancellation, refund, pricing exception, vulnerable customer or request involving sensitive informationDo not send a substantive answer. Route the case to a named person with the source message and a short summary.

This is deliberately conservative. Most small firms do not have thousands of identical enquiries that justify broad autonomy on day one. They do have a steady stream of repeat questions that AI can sort or draft faster, while keeping the final decision with somebody who understands the business.

The boundary can change after evidence from real work. Do not promote a category because a demonstration looked smooth. Promote it when the cases are consistent, the source information is reliable, the failure is containable and the business can see what went out.

What counts as a safe automatic reply?

Start with a message whose purpose is narrow: “We received your enquiry and a member of the team will reply by the next working day.” The promise about timing must match what your team can actually deliver. If you cannot reliably meet it, change the template.

A status reply needs more care. It should read from a trusted record that belongs to the verified customer. It must not infer a delivery date from a vague note, reveal a different person’s order or improvise a remedy when the record is missing. Missing or conflicting data should send the case to review.

Keep the send rule outside the model where possible. The workflow can check an approved enquiry type, an authenticated customer, a current source record and a permitted template. Only when every condition passes does it let the message leave. A model can help classify or draft, but the final permission can be enforced by ordinary software.

For an initial run, I would limit the number of automatic replies per hour and keep them within working hours. That gives the owner a chance to spot a poor pattern before it reaches a large group. The correct limit depends on the volume and consequence of the work; there is no universal number.

What should a person check before approving a draft?

An approval button is useful only if the reviewer has enough context to make a decision. Show the original enquiry, the proposed recipient, the draft, the records used to prepare it and any uncertainty the system found. The reviewer should be able to edit, reject or hand the case to another colleague.

I would ask the reviewer to check five things:

  1. Recipient: is this going to the right person and address?
  2. Facts: are the order, date, product, policy and any quoted amount supported by a current source?
  3. Promise: does the message commit the business to an action or deadline it can honour?
  4. Rights and access: does it preserve a customer’s route to cancel, complain or speak to a person?
  5. Data: does it reveal only information this recipient should receive?

Do not make the reviewer hunt through five systems while the AI has already produced an attractive answer. Put the evidence beside the draft. If the source is unclear, the approval choice should be “check manually”, not “send anyway”.

The CMA says businesses should consider telling customers when they are dealing with an AI agent if that fact could affect their decision. Avoid making an automated exchange look like a personal reply from a named colleague. Be plain about what the system does and keep an easy route to a person.

What customer data may the agent read?

Only the information needed for the defined task. The Information Commissioner’s Office discussion of agentic AI applies the familiar data minimisation principle: an agent should not receive information because it might prove useful later. Set its purpose, choose the records required for that purpose and limit its access accordingly.

An enquiry assistant that checks an order status may need that order’s status and the customer’s verified contact details. It does not need the entire CRM export, internal notes on unrelated accounts or a shared mailbox covering finance and HR. Separate access also makes a mistake easier to contain.

The same rule applies to what gets stored. Keep an appropriate record of the source, draft, approval and final message so the business can investigate a complaint. Limit access to that record, set a retention period that fits the business’s obligations and avoid storing unnecessary personal information in model prompts or logs.

How would I test it before a live send?

Take a sample of real enquiry types and remove personal details where a test does not need them. Include the awkward cases: missing order number, two customers with similar names, an expired offer, a complaint dressed as a routine question and a request for someone else’s information. Add a case where the policy changed yesterday.

For each case, write down the expected lane before running the agent. Then compare its proposed lane, draft and cited source with that answer. Count both wrong drafts and wrong routing. A system that drafts a clumsy sentence is fixable; a system that sends a refund promise from the wrong lane has a much more serious control failure.

Run the first live cases with every send approved. Record how often the reviewer changes the draft, which facts need correction and how much time the complete process takes. Include the review time. If approval takes longer than writing the answer from scratch, improve the workflow or stop the experiment. My Friday afternoon workflow test explains how to include exceptions in that first trial.

When a category performs consistently, test the exact automatic-send rule in a small live window. Inspect the sent messages, not just the drafts. A change to the model, prompt, customer data, template or connected system is a reason to review the rule again.

Who watches it after launch?

Give the workflow a named owner who can see sent messages, corrections, complaints and failed handovers. My AI ownership guide sets out the broader role. For customer email, the owner should also know how to pause sends immediately and how the team will handle the queue while the automation is off.

Watch a few measures that answer real questions: the share of enquiries routed to each lane, the proportion of drafts changed by a reviewer, the number of incorrect or disputed replies, and the time from enquiry to a useful answer. Review a sample of automatic sends, including quiet days with no complaints. Silence alone does not prove the replies were right.

The CMA advises businesses to correct a poorly performing agent quickly, especially when its outputs reach many people or vulnerable customers. Build the stop route before launch: disable the send permission, preserve the queue, identify affected customers and let people take over. Test that route once. A procedure that exists only in a document is hard to trust during a busy morning.

Where should a UK SME start?

Pick one repeated enquiry category. Write down what the agent may read, what it may draft, which messages need approval and which must be handed over. Test difficult examples, approve every initial send and measure the complete task. Only then consider a bounded automatic reply.

If you want help designing that boundary for your own workflow, my AI automation service covers the process, access and checks. You can get in touch with the enquiry category and the current hand-off process. That is enough to start a useful conversation.

Customer Service

Recommended Reads