InsightsGuardrails

Four things an AI agent should never do without asking

An agent that acts on your behalf is a new member of staff. The same rules apply: clear limits, supervision, and a record of everything it did.

5 min readBenjamin Yang

The useful thing about an AI agent is that it acts. It books the appointment, sends the invoice, replies to the enquiry. The dangerous thing about an AI agent is exactly the same. So before we build one we agree with the business, in writing, what it may do on its own and what it must stop and ask about. The list is shorter than people expect, and it is nearly always the same four things.

Move money

Taking a payment a customer has agreed to is fine. Issuing a refund, writing off an invoice, changing a price, or cancelling a paid plan is not, until a person has said yes. We set a threshold with the owner (often it is zero for refunds and a few hundred pounds for anything else) and the agent queues anything above it for approval with a one-line summary of why. The approval takes ten seconds. The absence of it is how a small mistake becomes a large one.

Pretend to be a person

Every agent we build says what it is in the first sentence, on the phone and in writing. Customers do not mind talking to an AI assistant. They mind being deceived about it. Beyond the ethics, there is a practical reason: a customer who knows they are talking to an agent asks for a person when they need one, and that hand-off is the safety valve for everything the agent was not built to handle. An agent that hides what it is removes the valve.

Keep going when it is out of its depth

Agents are confident. That is a feature when the rules are clear and a fault when they are not. So we define the edges: a caller in distress, a complaint, a question about medical or legal or financial advice, anything the agent has been asked twice and still cannot resolve. At the edge it stops, says it is passing this to a colleague, and does so, with a summary attached so the person does not start from scratch. Knowing when to stop is a large part of what makes an agent trustworthy.

Do anything you cannot see afterwards

Every action is logged: what came in, what the agent decided, what it did, and why. The owner gets a weekly brief in plain English, and anyone with access can pull up any conversation. This is not paranoia. It is the same standard you would apply to a new employee handling customers and money, and it is what makes it possible to tune the agent when it gets something slightly wrong rather than discovering the pattern three months later.

Where the rules come from

Partly from good practice, partly from regulation. GDPR shapes what data the agent may hold and for how long, and where it is hosted (for our clients, in the UK). FCA rules shape what an agent may say about a financial product. CQC and GDC expectations shape what a clinic's agent may and may not discuss with a patient. We build the relevant ones in by default and flag the rest, because the business is responsible for what its agent does, and it should know exactly what that is.

None of this slows the agent down in the ninety-five per cent of cases where the answer is obvious. It just means the other five per cent land with a person, which is where they belong.

Share post

More insights

AI roadmap

Stop buying AI tools before you know the job

The average small business we meet is paying for four AI subscriptions and using one. The problem is not the tools. It is buying them without a job in mind.

5 min readBenjamin Yang

Talk to us about your Tuesday.

A free 45-minute call. You tell us where the week goes; we tell you honestly whether an agent would help and roughly what it would save.

Book a call