The Maintenance Vacuum
Before another senior hire, find out which work actually left when the junior roles did
If you’ve cut junior and mid-level roles and the migration still won’t move, that distinction matters. Another senior can add capacity. But if the hire inherits the same unattended queue, the staffing plan needs to account for that work before promising more architecture or strategy.
Read past the ticket title
A dependency upgrade that changes how stored data is interpreted may warrant your most experienced engineer. Implementing a tested, reversible version bump may be a good assignment for someone still learning the service. The same ticket can contain both kinds of work. The risk and the available support determine who can own each part.
Google’s SRE book distinguishes repetitive operational work from engineering that leaves a lasting improvement. Its term toil describes work that tends to be manual, repetitive and automatable, with no lasting improvement to the service. Cleaning up an entire alerting configuration can be grubby work and still produce lasting value. A maintenance label tells you very little about which kind of effort you’re buying.
Google also cautions that recurring work can need expert judgment and still be toil if a redesign could remove it. For a staffing plan, that means asking who can handle the task safely today and what would stop it coming back. A senior may be needed for both, but handling the next failure and removing its cause are separate assignments.
The staffing implication is conditional: if a task survives the removal of its owner, someone else must do it or it must wait. When it falls to a senior, that engineer has less time for other assignments. Whether the trade was worthwhile depends on what got displaced and what the business saved. A smaller payroll can be a real saving; a delayed migration can be an acceptable trade.
The maintenance vacuum opens when the plan counts the saving but leaves the surviving work unassigned, then expects the remaining engineers to deliver the old roadmap anyway.
What it takes to remove the work
In Google’s account of automating datacenter network repairs, engineers had to check whether traffic could safely move off a failed device, move it away, and restore it after a repair. As the network grew, the volume of problems began to overwhelm the engineering staff.
The team built a system that assessed the risk, moved traffic, and handed physical repairs to technicians when needed. The authors report that it freed engineers to work on Jupiter, the next generation of the network. They describe the work that changed and where the recovered capacity went.
The first workflow also introduced problems. Its interface could disagree with the switch’s actual state. In one episode, a technician started concurrent operations to move traffic away from devices awaiting repair, causing congestion and user-visible packet loss. The software allowed those concurrent operations; the authors say they hadn’t tested it with new technicians. They redesigned later workflows so equipment was ready for repair before the technician arrived.
The repair software needed maintenance of its own. Parts shortages for the newer equipment prolonged the life of the older equipment, so engineers still had to improve the automation supporting it.
That account concerns network operations, with no comparison of junior versus senior hiring. The useful inference for a staffing plan is about the assignment: reducing a recurring queue takes engineering work of its own. An experienced hire can help do that work, provided the hire has room to change the process rather than only keep it running.
The same question applies when AI or an outside team is supposed to absorb maintenance. Follow a completed upgrade through testing, review and release, including rework or a rollback. If a tool produces the patch but a senior still investigates failures and gets it into production, count that remaining effort. A supplier that owns review and release may remove more of that effort, though your coordination and escalations still count. A smaller patch-writing bill alone won’t tell you how much capacity your own team recovered.
Keep the learning visible
Then ask what the maintenance was teaching somebody.
Working through an unfamiliar dependency change with an experienced reviewer can teach an engineer how compatibility gets tested and when a rollout should stop. Repeating a procedure they’ve already mastered offers much less of that. The part worth preserving is supervised responsibility, followed by harder work as their judgment develops.
If you automate the routine steps, keep opportunities to investigate exceptions, improve the tests, and explain release decisions. If you outsource the whole workflow, decide where engineers inside the company will get that experience instead. Removing repetitive work and preserving a learning path are separate decisions.
That learning needs time from the people doing the reviewing. A junior hire needs a bounded assignment, feedback and someone available when the work exceeds their experience. If reviews are already the bottleneck, adding someone who needs the same review can lengthen the queue. Hiring someone able to share those decisions may offer more relief.
Hire against the work that’s left
Take a sample of recent maintenance work with the engineers who did it, then add work still waiting for an owner. If roles were cut, compare what those people handled with what remains: some tasks may have been deliberately stopped, others transferred or deferred. Completed tickets alone can’t show you the work nobody picked up.
Record who handled each part, how much attention it required, and what waited while they did it. Keep active work separate from a ticket’s age: an upgrade waiting for review needs a different response from one consuming a senior’s attention with compatibility failures. Include supervision and rework, whether the first draft came from a colleague, a supplier or AI.
If the constraint is genuinely difficult maintenance, another experienced engineer can be the right hire. If it’s repeatable work with a clear way to teach and review it, a junior or mid-level role may fit.
When the queue keeps refilling with the same avoidable task, compare the work it consumes with the cost of removing its cause and maintaining the replacement. The existing team may be able to make that change if you explicitly take something else off its plate. For an infrequent task that’s safe to handle manually, assigning it to an existing engineer can be cheaper than building another system to own.
The next staffing conversation shouldn’t open with whether your senior engineers are good enough. Open it with what they spent Tuesday on.
Want this looked at in your business?
Start with the rough map: thirty minutes, owner to owner, and a written report on where AI pays off for you and what it's worth. It's free, and if you don't need us the report says so.
Get your rough map, freeNot ready to talk? Stay sharp anyway.
We send insights like this to technical leaders every week or two. The thinking we bring to our engagements, no fluff, no spam.
You're in. Check your inbox to confirm, and for what we sent.
