What Your AI Receptionist Is Allowed to Promise
A routine-sounding call can ask for a decision your business hasn't made.
A confident yes could leave you choosing between funding a repair you never approved and disappointing a customer who called your number and believed the answer. A refusal to discuss anything involving money would go too far the other way. Reading an approved callout fee can be a useful part of answering the phone.
That distinction matters more than how human the voice sounds. If routine questions are interrupting your staff, answering them automatically has a clear purpose. But the same caller can move from asking your opening hours to asking for a warranty exception without hanging up. The agent needs to recognize when the request exceeds its authority.
The wrong answer can sound like company policy
In November 2022, Jake Moffatt asked Air Canada’s website chatbot about bereavement fares. It said they could apply for a reduced fare after travelling. Moffatt booked flights relying on that advice, then learned that the airline’s policy prohibited retroactive applications.
In February 2024, the British Columbia Civil Resolution Tribunal found Air Canada liable for inaccurate advice Moffatt reasonably relied on. It ordered the airline to pay C$812.02, including interest and tribunal fees. The chatbot had even linked to the policy page that contradicted its answer. The tribunal rejected the idea that Moffatt should have known which part of Air Canada’s website to trust.
That detail matters when a vendor offers to load your policy documents. Having the right policy available and giving the right answer are separate things. The decision doesn’t establish how Air Canada’s chatbot produced the error; it establishes what the customer received and relied on.
In April 2025, Cursor’s email support bot explained unexpected logouts as a new login policy, according to Fortune. The company had made no such policy. Cofounder Michael Truell later said his team had fixed the session bug, refunded the affected user, and begun clearly labeling AI support replies. The support answer had turned a software fault into an apparent business decision.
Both systems delivered an answer someone could act on, through a channel the business provided. These were text exchanges, and neither case measures how often a voice receptionist will fail. The operational inference is narrower: a fluent answer about your business can carry authority that its accuracy hasn’t earned.
What “let me check that” has to mean
A receptionist who says “let me check that” buys time to consult a record or find someone authorized to decide. Saying the words without doing either would be useless, whether the speaker is a person or a model.
Vendors do provide handoff tools. RingCentral’s October 2025 release notes describe transferring callers to a human and showing the receiving person a summary in its desktop app. A business may already have the features it needs. It still has to specify when to use them and who can answer the question on the other end.
For the warranty call, draw an escalation map around the promise, rather than banning the whole subject. For each answer the agent may give, name the record or approval that supports it. For an answer it must defer, name who receives the request and what the caller can expect if that person is unavailable.
An agent can be authorized to explain the published warranty terms without deciding a claim. If your service record already contains approval for the repair in question, relaying that decision can also fall within its remit. But if coverage still depends on an inspection, the agent should collect the request for review and leave coverage unresolved. An exception to the published terms belongs with whoever can approve it.
The same distinction applies to bookings. If the system has actually reserved an available appointment, it can confirm that booking. If it has only taken the caller’s preferred time, it must describe that as a request. This is where an otherwise helpful receptionist can create work for your staff: the caller hears a commitment while your team receives an inquiry.
Write those boundaries with the person who currently handles the exceptions. A vendor can help translate them into routing rules, but a generic instruction to be helpful doesn’t specify who may waive your fees. A generic instruction to escalate difficult calls leaves the agent to decide what counts as difficult.
If the map exists only in the agent’s instructions, the model still has to recognize when to hand off. It could treat a claim requiring approval as a general question and keep talking. A booking system that rejects unavailable appointments can prevent an invalid reservation, yet the voice could still claim success. The boundary has to hold in what the caller hears as well as in what the software does.
Test the call nobody can take
Use the warranty question in an after-hours test call, with the policy and service record open beside you. Move from asking about the published terms to asking whether your particular repair is free. Try both a claim awaiting inspection and a job whose coverage is already approved. That tests whether the agent invents a promise when facts are missing, but also whether it needlessly defers an answer it’s authorized to give.
Then leave the transfer unanswered. Follow the request to the queue where it should land and check what the caller was told. The summary should preserve a request for coverage as an unresolved question, rather than turn it into an approved warranty job. The receiving person needs the caller’s details and authority to decide, or a route to someone who has it. A notification in an unattended inbox leaves the work undone.
Test an unavailable appointment too, and a reservation that fails to complete. Compare what the caller hears with what the calendar actually contains. A receptionist that claims a booking while leaving only a request has failed this test, even if its software correctly refused the reservation.
A successful rehearsal covers the calls you tried; changed policies and unfamiliar requests still need review. These tests can expose a broken boundary. They don’t establish how often an agent will cross it in live use, or how much staff time it will save.
If after-hours callers can wait, voicemail with a reliable callback process may be enough. If callers need immediate decisions that a person with the right records and authority can make, staffing the line or using a human answering service is an alternative. Apply the same test to that service: an operator can take the request, but settling it still requires the relevant records and permission. Where coverage requires an inspection, whoever answers has to leave that decision pending.
For the Saturday-night warranty caller, a promise of a Monday callback is itself a commitment. Before the agent makes it, decide whose queue will hold the request and whether that person can actually make the call.
Want this looked at in your business?
Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.
Get your rough map, freeNot ready to talk? Stay sharp anyway.
We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.
You're in. Check your inbox to confirm, and for what we sent.
