Back to Insights

The Paved Demo

A good demo shows a selected case. Before you buy, specify the work the tool must finish and what should happen when it can't.

6 min readBy The Bushido Collective
AI StrategySmall BusinessProcurement
Share:LinkedInX
If an AI receptionist tells a customer their appointment is booked, but nothing reaches your staff’s calendar, the customer heard a promise your team can’t keep. A demo can stop at the convincing reply. Your evaluation has to follow the booking.

When the vendor picks the input, prepares the records, and chooses where to end the demonstration, you have evidence about that setup. You still need to find out what happens with duplicate customer names, missing details, or a calendar that refuses the booking.

Call it the paved demo: a smooth result on a route selected in advance. Preparation is legitimate. It lets a seller explain a product without spending the meeting on configuration. The mistake is treating success on that route as evidence that the rest of your work is covered.

What supports the promise?

The FTC’s case against DoNotPay illustrates how specific a missing test can be. The agency alleged that the company advertised a substitute for a human lawyer without testing whether its output performed at that level. The final order, approved in January 2025, prohibited DoNotPay from advertising lawyer-equivalent performance without sufficient evidence.

In March 2024, the SEC announced settlements with investment advisers Delphia and Global Predictions over charges that included false and misleading statements about their use of AI. The firms agreed to $400,000 in total civil penalties without admitting or denying the SEC’s findings.

Those cases don’t measure how common misleading demos are. The lesson we draw is narrower: ask what supports the exact promise you’re buying. A label on a product and a successful example answer different questions from whether it can finish your work.

You can evaluate that promise without reading the code. Your part starts with specification: describe the result you need and the conditions under which the tool should stop. The person who handles bookings already knows what information must arrive and which mistakes would cause trouble. Bring them into the evaluation.

Follow the work past the answer

For an appointment-booking tool, define success at the calendar: the correct service, time, and customer details are recorded once, and the customer receives a matching confirmation. Have the person who normally handles bookings open that appointment from their own account. A confirmation shown only on the vendor’s screen leaves the handoff untested.

Start with an ordinary request your business accepts, under the service and availability rules you’ve supplied. Then try a customer name that matches more than one record, or a request outside your service area. Agree beforehand whether the tool should ask for missing information, decline, or refer the request to staff. In those cases, a pause can be the correct result.

Next, ask the vendor to make the calendar connection fail in a safe test environment. If the calendar refuses a booking, watch whether the tool tells the customer it remains incomplete. If the calendar never sends a reply, the outcome may be uncertain: it could already contain the appointment. Retrying without checking could book the customer twice.

The staff handoff should carry that uncertainty along with the customer’s request and contact details. Have someone on your team pick it up from the tool they normally use, establish whether a booking exists, and finish or correct it. Check what the customer is told next. If someone must copy details between screens, count that as part of the work you are buying.

Use fictional test records that preserve the awkward structure of your work without exposing customer identities or confidential information. A sales meeting doesn’t justify uploading a production export. If a later trial genuinely requires real records, settle access, retention, model-training use, and deletion terms before sharing them, with the appropriate privacy and security review.

Give the vendor the requirements in advance. Connecting your calendar can require permissions, configuration, and work by your own staff. Ask what must be prepared, who will do it, and what it costs. If records need cleaning or fields need mapping between systems, put that work in the proposal.

A follow-up for that preparation is reasonable. What matters is whether it leads to the agreed test, with the preparation visible. If a requirement remains untested, leave it marked unproven. You don’t need to decide whether the salesperson is dishonest to withhold acceptance of a capability you haven’t seen.

What a successful test earns

After setup, run additional requests you select from the agreed scope, without rehearsing those exact inputs. Keep the business rules fixed; a surprise requirement would test the vendor’s guess rather than the promised capability. Record any help needed during the run. A result that needed a vendor employee to repair the record includes that person’s work. If the promise is unattended booking, rerun the case without that intervention.

One successful difficult case establishes that observed result, under those conditions. Keep the misses and staff corrections in the record alongside successes, including when a fix turns a failure into a pass. Because you chose these cases deliberately, their mix needn’t resemble everyday bookings. A small trial can expose a missing handoff; it can’t establish a dependable error rate.

Match the next commitment to the evidence. If the tool offers a staff-approval mode, test the hold itself: have staff reject a proposed booking and check that neither an appointment nor a confirmation is created. Giving a tool permission to confirm bookings on its own calls for evidence about failures and recovery as well as correct answers. Keep that permission limited while those questions remain open.

The tool may still be worth buying if a person has to review every booking. Have that person do the review during the trial, and compare the total work with handling similar requests through your current process. Include checking the calendar, chasing missing information, and correcting errors. If you’re buying time back and that work consumes the saving, keeping the current process is a reasonable outcome of the trial.

The demo needs to reach the moment your team has enough information to take over, even when the AI has nothing useful left to say.

Want this looked at in your business?

Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.

Get your rough map, free

Not ready to talk? Stay sharp anyway.

We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.

Keep reading

Share:LinkedInX