Friction Was a Feature
When AI lowers the cost of building, a useful experiment still needs a question it can answer and a decision someone will own.
Dax Raad captured the uncomfortable part in a February 14 post: “ideas being expensive to implement was actually helping.”
Cost can force a sponsor to explain why a proposal deserves resources that another idea won’t get. If AI makes the implementation cheap enough to bypass that decision, the scrutiny can disappear with the price.
But cost was a crude filter. It could make an idea compete for attention without telling you whether customers would value it. Restoring the old approval burden would bring back that limitation too.
The idea the filter missed
In 2012, a proposed change to Bing’s ad headlines sat at low priority for more than six months. An engineer saw that it would be cheap to implement and ran a controlled experiment comparing versions.
In their 2017 Harvard Business Review account, Ron Kohavi and Stefan Thomke describe revenue rising enough to trigger a “too good to be true” alert. Such alerts usually pointed to a bug. This time, analysis confirmed a 12% revenue increase without hurting the key user-experience metrics.
The result belongs to that Bing experiment. It predates today’s AI coding tools and gives us no estimate of their value. It does give us a reason to distrust the idea that a feature must persuade leadership of its worth before anyone tests it: the prioritization process had passed over a winner.
So building more, measuring the results and killing the losers is a serious answer to cheaper code. If the experiment is contained, the results are interpretable and removal is affordable, lower implementation cost lets you test a possibility without committing to support it indefinitely.
Bing had invested in the work around those tests. In a 2013 paper on large-scale experimentation, its researchers describe how repeated trials can create false positives: apparent improvements produced by chance even when a feature does nothing. They adjusted statistical thresholds and reran the final candidate to check the result, rather than treating the best-looking trial as the answer.
That effort extended to checking for tests that affected one another. The researchers also report avoiding deployment of features with negative results despite stakeholders’ early enthusiasm. Their system had to support finding out that an attractive idea was harmful, then acting on that answer.
Budget for an answer
The distinction we draw is between a cheap implementation and a cheap experiment. The latter includes getting evidence you can use and being able to stop. A polished demo can show that a feature works while leaving both costs unexamined.
Take a proposed onboarding change. If its purpose is to help new accounts finish setup, clicks on the new screen won’t answer the whole question. Suppose the new flow lets people skip a required import: completion could rise while more accounts reach their first task without the data they need. Define what successful setup leaves the customer able to do before comparing completion rates.
A customer walkthrough can expose confusing steps; a controlled production comparison can test whether completion improves. For that comparison, randomly assign new accounts to the old or new flow and measure the same outcome in both. If the trial group also gets extra help from your team, you’ve tested the screen plus the help. With few new accounts, a small effect may remain uncertain however quickly the team can generate another version.
That uncertainty changes what is worth building. A disposable prototype may be enough to investigate the confusing step. A customer-facing release brings another question: if people store data in the feature or build a process around it, what happens to their work when you remove it? Deleting the code leaves that question open.
Before implementation, the sponsor and engineer can agree on the question and what evidence would justify a wider rollout. That can fit in a ticket; it doesn’t require pretending to know whether the idea will work. Scale the approval to what the test puts at risk. Asking a throwaway prototype to justify the support burden of a permanent feature would rebuild the old filter.
This can still mean declining a plausible-looking initiative in front of people who could build it. The reason might be that you can’t get useful evidence from it, or that its support obligations would displace more important work. Those are arguments the sponsor can challenge. They can offer a smaller test, change the scope or defend the tradeoff.
And if you authorize the experiment, agree who will close it. A negative result needs a keep-or-remove decision; an inconclusive result needs a decision about whether more evidence is worth pursuing. Put the sponsor’s name beside that decision before the demo, especially when the sponsor is you.
Want this looked at in your business?
Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.
Get your rough map, freeNot ready to talk? Stay sharp anyway.
We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.
You're in. Check your inbox to confirm, and for what we sent.
