Somebody Still Presses Send
What human approval contributes, and when it limits the return on your AI investment.
Less writing can be worth paying for. But if the purchase was supposed to mean faster service, the renewal conversation needs more than a count of AI-assisted tickets. It needs to follow the work through to a correctly resolved problem.
The Supervision Default
AI proposes, a person approves. At launch, that’s a reasonable response to uncertainty: someone should notice an error before a customer has to live with it. It becomes the supervision default when the rule survives without anyone asking what the reviewer must catch, whether they can catch it, or which work still needs them.
The person pressing send may also be the person who used to write the whole reply. If assistance leaves them less writing to do, it can free up time to check and resolve cases. Keeping their approval tells you little about whether the workflow got faster.
If drafts already arrive faster than a fully occupied reviewer can clear them, making more drafts lengthens the queue. Completed work increases only if review can process more cases or fewer cases require it. Under those conditions, another improvement in writing speed won’t produce a matching improvement in service.
Follow requests from arrival through final resolution. Separate writing and checking time from time spent waiting, and locate the wait before deciding which step to remove. A reviewer who is available but can’t see the account history needs access to that history; removing their approval does nothing to supply it.
Keep the Boundary Revisable
There is good evidence that supervised AI can improve completed work. In Generative AI at Work, published in 2025, researchers studied customer-support agents serving a single business-software company. They estimated an average 15% increase in issues resolved per hour after access to AI assistance. The human agents remained responsible for the conversations and could ignore or edit the suggestions.
Gains varied substantially across workers: less experienced agents benefited more, while the most skilled saw small declines in quality. The study didn’t compare this setup with autonomous support. It does establish that retaining a human’s final authority can coexist with a measurable productivity gain. The supervised workflow deserves measurement before anyone declares it wasted capacity.
Autonomous service needs room for correction, too. In May 2025, Bloomberg reported that Klarna CEO Sebastian Siemiatkowski said his AI-fueled customer-service cost-cutting had gone too far. He was planning a recruitment drive to ensure customers could speak to a person.
Yet in an interview with Big Technology the following week, he said Klarna was expanding the work its AI could handle. The company was also hiring people for higher-end conversations it had previously outsourced. Both changes were possible at once.
A hiring plan and a claim of expanded AI use leave the service outcome unresolved. Our reading is that the autonomy boundary has to remain revisable in either direction. Increasing human support in one part of a service doesn’t require retreating from AI everywhere.
What Did Review Catch?
Start with a defined category, such as payment-status questions, and sample cases from it rather than selecting only the replies reviewers edited. Keep the original draft alongside the final reply and the eventual resolution. Count a corrected policy claim separately from a rewritten greeting. If a reviewer stops an incorrect refund, record the decision even if no message goes out.
An unchanged reply is ambiguous. It could mean the reviewer confirmed it was correct, or that the reviewer missed the same error as the model. Have someone assess both versions in the sample against the account record and policy, rather than treating the human’s final version as the answer key. A count of edits alone cannot tell you whether approval protects the customer.
If reviewers keep catching consequential errors, preserve that check while addressing the cause. If a model used an obsolete refund rule, update its source and then check whether it answers the same cases correctly. Moving the same poorly informed decision into an unattended process would only make it harder to interrupt.
Where a narrowly defined category holds up under those checks, it becomes a candidate for fewer approvals. You can assess drafts before changing what reaches customers; learning doesn’t require switching off review for the whole operation.
A clean sample still leaves rare, severe failures unresolved. If an error could disclose another customer’s records, a later correction cannot undo it. Keep high-stakes or unfamiliar work supervised, along with anything that requires human authorization by law or contract.
For a payment-status request from an authenticated customer, a template and a lookup may be enough. The system would return that customer’s recorded transaction status without initiating a refund. An unclear account match or disputed charge would go to a person. Choose the simpler mechanism when the work supports it.
The handoff needs testing as much as the reply. If the system confidently labels a dispute as a status request, it can answer the narrow question correctly while leaving the customer’s actual problem untouched. Try such cases before reducing approvals, and confirm they reach someone able to resolve them.
Any live trial still needs a bounded downside and errors that can be detected and repaired before they cause material harm. Compare it with supervised requests from the same category; automated status lookups and human-handled disputes are different work. Measure correctly resolved requests and reopened cases, alongside waiting time and the time spent handling exceptions.
Include tool fees and the work of maintaining the system or repairing mistakes before calling it a saving. Otherwise the apparent saving may be work you moved to someone else’s queue.
Want this looked at in your business?
Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.
Get your rough map, freeNot ready to talk? Stay sharp anyway.
We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.
You're in. Check your inbox to confirm, and for what we sent.
