The Firefighter Premium
Making preventive work visible without turning it into another rescue story.
When a review celebrates the rescue but records preventive work as routine upkeep, it pays a firefighter premium: more recognition for handling a crisis than for removing its cause. The recovery deserves credit, and the person restoring service may have had no part in creating the fault. The distortion comes when an engineer needs a failure before the organization can recognize the value of fixing it.
Nelson Repenning and John Sterman documented this incentive in their 2001 research on process improvement. One project leader described delivering programs under duress as a route to advancement: “Our [company] culture rewards the heroes.” Their research combined industrial and product-development case studies with models of how work pressure crowds out improvement. It supplies a mechanism and reported examples, rather than a measure of how frequently software companies reward firefighters today.
In their model, postponing maintenance or training frees time to meet today’s target. The damage arrives later, as equipment becomes less reliable or people lack the preparation to handle the work. The resulting trouble draws still more time away from improvement. The authors describe rewards for rescuing troubled work reinforcing the same short-term behavior that helped create the problem.
Our inference for an engineering budget depends on what changes between the routine request and the escalation. New demand or newly available funds can justify a different decision. But if the facts and resources stay the same, and calling the request an emergency is what gets it approved, the requester has a reason to make the next request sound urgent.
That puts a limit on the usual advice to tell a better story about your work. A frightening account of the outage you supposedly prevented can win attention without establishing that the work was necessary. A clean quarter is compatible with good preparation, excessive spending, or luck.
The useful record starts before the outcome is known, in the work ticket or decision record the team already uses. It connects the evidence of a failure risk to the cost of addressing it, including reasons to leave the system alone. A reviewer still has to check that evidence; a polished write-up can exaggerate a risk as easily as a dramatic presentation.
Suppose a load test shows checkout requests timing out when an internal report runs against the same database. A planned sales campaign would put more demand on checkout. Calling the proposed work a database improvement hides the choice: should the campaign depend on those competing workloads? Separating reporting from checkout might be justified. Rescheduling the report could be cheaper and sufficient.
Before choosing, preserve the test that exposed the problem and the campaign forecast that makes it relevant. Include the basis for that forecast. A timeout at a load the business has no reason to expect gives a weak basis for delaying feature work. Product leadership can question the forecast and agree what checkout must handle before the team tests a fix. A less demanding retest could make an ineffective change look successful.
The cheaper remedy needs the same scrutiny. Rescheduling is sufficient only if the report can run reliably outside the busy period and still reach its users when they need it. If the team proposes a separate reporting database, its estimate has to include keeping that copy current and operating it. Put those costs beside the feature work each option would displace.
After the change, repeat the test under the same checkout load and data conditions. For the scheduling option, check that the report actually runs outside the busy period and finishes when its users need it. If checkout meets the agreed requirement and reporting still works, you have evidence that the intervention addressed the demonstrated conflict.
Keep the conditions beside the result. A successful test doesn’t establish how much revenue would have been lost, or prove the system can survive workloads you haven’t tested. If the campaign later draws less traffic than forecast, record that too: the test still shows what changed, while the uneventful launch tells you less about whether the work was necessary.
The review needs the actual cost, including any feature work displaced. In Google’s account of reliability tradeoffs, Marc Alvidrez explicitly includes the opportunity cost of assigning engineers to reduce risk instead of building features. If the existing system meets the business’s needs and the feared workload remains speculative, deferring a redesign can be a defensible decision. A credible record must leave room for that answer.
Even demonstrated results need someone willing to protect the work. Repenning and Sterman report that direct maintenance costs fell by an average of 20% at Du Pont plants implementing an improvement program by the end of 1993. They also report that corporate downsizing weakened the program’s ability to reinvest savings in further improvement. The case describes a broader maintenance program, so its result cannot be assigned to communication alone. It also shows why documenting value cannot guarantee continued funding.
You can change what gets a hearing in the reviews you control. If an engineer reproduces the checkout failure and shows that rescheduling meets the agreed requirements, give that decision the attention you’d give an incident recovery. Ask to see the test and the tradeoff while there’s still no incident to discuss.
Want this looked at in your business?
Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.
Get your rough map, freeNot ready to talk? Stay sharp anyway.
We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.
You're in. Check your inbox to confirm, and for what we sent.
