The In-House Mirage
Owning the code doesn't settle who has the experience to build and run it.
Before turning that into a judgment on the team, name the work that remains. Can the system find the right information? Can someone tell when its answer is wrong? Who has time to fix it? Keeping the code inside the company settles none of those questions.
There’s a tempting shortcut in MIT Project NANDA’s preliminary 2025 report. It reports deployment rates of roughly two-thirds for external partnerships using customized tools designed to learn from feedback, compared with one-third for internal builds. That comparison comes from an interview sample of 52 organizations, with self-reported outcomes and varying definitions of success.
The authors explicitly warn that the relationship may reflect differences in organizational capability rather than the delivery approach alone. One possible explanation is that a company buying a tool for a well-understood task faces a different problem from one building something no vendor offers. The report also lists hybrid teams, internal developers working with an external vendor, as having insufficient data to quantify. It cannot tell you the odds of hiring a specialist to work alongside your engineers.
An internal plan still has to account for the work it displaces. If the same engineers are maintaining your product, the AI build consumes time that already had a job. Their salaries are paid either way; the postponed work still has a cost. Add unfamiliar operating problems, and the project can demand more than the team has available even while every line of code stays yours.
An outside plan has to reserve your people’s time too. Someone must explain the exceptions and review whether the system handles them correctly. If nobody inside can do that work, an additional builder can still be left waiting for decisions.
Think about framing a house. You wouldn’t hand your sharpest generalist a framing nailer and a stack of tutorials just because they’re available. Mistakes can go inside the walls, where nobody finds them until the roof starts to sag. You’d want someone who knows how to check the frame before you close it up. An experienced builder already on your payroll qualifies.
Your own people also hold knowledge an outsider has to acquire. But the person who understands the business process may be a support lead or an accountant, rather than the engineer assigned to automate it. In RAND’s 2024 study of AI project failures, practitioners described models optimized for the wrong business metrics and developers needing subject-matter experts to judge which data were reliable. Those interviews concerned machine-learning development and excluded projects that simply used pretrained language models.
Consider a proposed refund assistant. Suppose the standard policy would deny a refund, but an account note records an approved exception. If the assistant answers from the policy alone, better prose won’t repair the result.
The support lead has to confirm that the exception still applies before you treat a refund as the correct answer. A disagreement there needs a policy decision, whoever writes the code. Once the answer is agreed, follow the information: an inaccessible note calls for work on how the assistant retrieves records. A note that reaches the assistant but gets ignored calls for a correction to how it makes the decision. An outside specialist who never learns about those exceptions can build the wrong thing just as efficiently as your own team.
Test the case again without the approved exception; the standard policy should now apply. An assistant that grants every refund could pass the exception case while breaking your policy. If it can issue refunds rather than just draft replies, check the recorded transaction too. Keep uncertain cases with a person while you test the wider workflow.
A builder who has operated a comparable system can help design checks for a confidently wrong answer, or for an action that breaks the next step in the workflow. Ask them to show what failed in that system and how they detected it. Apply that question to an outside bidder as readily as to an internal lead. A demo that avoids your difficult cases leaves the main question unanswered.
If your team already has that experience and access to the people doing the work, keeping the build internal is a reasonable choice. Give it capacity instead of wedging the project between existing commitments.
An internal build can also be how your team acquires the experience. If the workflow is central to the business and will keep changing, learning to maintain it has value beyond this release. Budget for review and limit what the unproven system can do while that learning happens.
If the missing skill is narrower, such as how to test answers, a focused review may be enough. Your team still has to run those checks and handle the failures. And if an existing product can pass your workflow tests at an acceptable cost and under acceptable data terms, compare it before funding a custom version.
Bringing someone into your codebase remains an option when a missing skill justifies it. Retaining control takes more than possession of the finished files: your team needs access to the running system, tests it can rerun, and the ability to change or disable the behavior after the specialist leaves. Put the relevant rights and handover obligations in the agreement. Before accepting the work, have your own engineer change the refund rule, rerun the checks and disable the assistant without the specialist driving.
At the next demo, ask the person who’ll use it to bring a case they still can’t safely hand over. Have the builder work through what fails and what would need to change. That gives you a place to spend the next dollar, or a reason to keep it.
Want this looked at in your business?
Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.
Get your rough map, freeNot ready to talk? Stay sharp anyway.
We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.
You're in. Check your inbox to confirm, and for what we sent.
