Back to Insights

The Pause in the Board Update

Buying the tools is easier than reopening the rules. Some rules can go; others still protect your customers and money.

6 min readBy The Bushido Collective
AI StrategyAI TransformationLeadershipEnterprise AIDigital Transformation
Share:LinkedInX
If your AI update can show who uses the tools but can’t show what the business gets in return, a board question about results leaves you reaching for a pilot anecdote. Take customer support: an assistant can draft a reply quickly while the refund still waits in the same approval queue.

The drafting improvement may pay for itself. Serving the same customers to the same standard at a lower total cost is a real result, even if the approval queue stays. But quicker replies alone establish neither that saving nor a shorter wait for the refund. A customer waiting for their money has little reason to care how quickly the holding reply was written.

What Uber Reconsidered

At Davos on January 20, Uber CEO Dara Khosrowshahi described this problem inside Uber. He said having AI follow the company’s existing customer-service policies produced some success. The more promising approach, in his account, came from rebuilding around an underlying goal, such as how the customer feels at the end of the interaction, and allowing the agent to reason through it.

His assessment deserves the same scrutiny as your own board slide. Business Insider’s report gives no before-and-after service measurements, no list of the policies removed, and no account of the financial or safety limits retained. It supports asking whether the old process still fits. It leaves the results, and the boundaries around that experiment, unproven.

Those boundaries matter. Given only a goal of making customers happy, an agent that refunds every complainant could appear successful. The objective alone gives it no reason to reject that solution.

Why the Rules Are There

A customer-service handbook can put a refund ceiling beside a requirement to escalate every refund. They look alike on the page. One limits how much money can leave the business; the other determines who gets to act. Either could be protecting the business from a real loss.

A rule can be scar tissue: reasonable when it was written, still treated as load-bearing after the original constraint has gone. The work is finding out what it protects before deciding that automating the task makes it unnecessary.

Suppose a support rep must hand a refund request to another team because only that team can see the payment record. An AI that writes a better handoff note leaves the customer in the same queue. If an agent can securely inspect the relevant record and request an eligible refund, the handoff may become unnecessary.

But if the second team also provides an independent approval to prevent fraud, access to the record solves only part of the problem. That approval needs to remain, or its owner needs evidence that a replacement control covers the same risk. The ability to look up a payment doesn’t settle who may send money back.

The refund ceiling still has a job. So do identity checks and protection against issuing the same refund twice. Removing the handoff safely would require the transaction system to enforce those checks independently of the agent’s decision. Those checks bound what it can pay; they don’t prove it understood the complaint. When the evidence conflicts or the request falls outside its authority, a person still needs to take over.

There may also be a cheaper way to remove the queue. If record access is the only obstacle, give the rep authorized access. If eligibility follows fixed conditions in reliable records, ordinary software can check them. An agent may help interpret a customer’s explanation, but that benefit has to survive the cost of checking and correcting its decisions.

Buying a subscription leaves those decisions largely alone. Changing who can issue money requires the people responsible for service and financial loss to agree on a new boundary. That’s why examining the rulebook reaches further into the organization than procuring the tool.

Klarna’s More Complicated Lesson

Klarna’s public accounts show how difficult it is to judge the result even when the operation changes. In May 2025, CX Dive reported CEO Sebastian Siemiatkowski’s admission that an excessive focus on cost had produced lower-quality customer service. The company was piloting recruitment of human support staff. Its spokesperson also said the chatbot still handled two-thirds of customer inquiries.

Yet Klarna’s Q1 2025 release, published May 19, reported customer-service costs per transaction down 40% since Q1 2023 while maintaining customer satisfaction. Those are company-reported figures. The release doesn’t explain how the satisfaction claim fits with the CEO’s quality concern, or break it down by type of request.

That uncertainty prevents a tidy verdict. The public evidence establishes neither that Klarna abandoned AI nor that unchanged rules caused its problems. The statements could concern different customers or different definitions of quality. A good average can coexist with a badly served group of customers; whether that explains Klarna’s figures remains unknown.

Cost per transaction also answers a different question from cost per resolved support request. Neither measure, by itself, tells you what happens to a customer whose case needs a person after the chatbot has finished with it.

For the refund example, start the clock at the customer’s first contact and carry it through any handoff to the final decision and, if a refund is due, payment. Keep open requests visible rather than calculating speed only from cases that happened to finish. A repeat contact about the same unresolved refund belongs to that request, even if the chatbot records a new conversation.

Compare similar cases before and after the change: routine refunds with routine refunds, disputed ones with disputed ones. Include the model’s running costs and the human work, including review and corrections; show the implementation cost separately. Check wrongful refusals as well as incorrect refunds. If only the easiest cases reach the agent, report its result for that group rather than treating it as the performance of the whole service.

The comparison needs the simpler alternative, too. If a rep with the same authorized access can resolve those requests just as quickly, the AI version needs to justify any extra cost or risk. Beating the old queue alone leaves that question unanswered.

The Question Behind the Slide

A cautious pilot can be honest work. So can a useful assistant inside a process whose controls still earn their place. AI theater begins when evidence of tool use is presented as evidence of a result that nobody measured. A deleted rule earns no credit if it merely moves the cost to a customer or leaves someone else cleaning up.

At the next board update: what improved, what did it cost, and what evidence says the change was worth keeping?

Want this looked at in your business?

Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.

Get your rough map, free

Not ready to talk? Stay sharp anyway.

We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.

Keep reading

Share:LinkedInX