I recently returned from 10 days of PTO and had a one-on-one with my VP, who is also my direct manager. Our team consists of me, a senior infrastructure engineer, and three mid-level security engineers. He said we need more cross-training, especially on the infrastructure side, because the organization is effectively "dead in the water" when I'm away.
Management has said there will be no additional infrastructure headcount, so we're expected to make do with the current team and our MSP. I've raised the need for redundancy in my role several times over the past few years, both verbally and in writing, but nothing was done.
I already provide cross-training and have documented procedures and workflows. I'm willing to do more, but I don't think it's realistic to create an SOP for every possible troubleshooting scenario. Troubleshooting requires technical judgment and experience, not just a checklist.
I also keep getting contacted during PTO when something goes wrong, which makes it difficult to actually disconnect. My VP now wants more documentation and cross-training, but there doesn't seem to be any extra time or staffing to support that work.
Has anyone dealt with a similar single-person dependency? What helped reduce the workload, improve coverage, and make sure PTO was actually protected?
4 Answers
You’re right that nobody can write a troubleshooting guide for every possible failure, but useful documentation still has value. Focus on system ownership, dependencies, diagrams, access paths, common failure modes, escalation contacts, and the first checks someone should perform. Decision trees and flowcharts can be more useful than long narrative documents.
After every major incident, hold a short review and capture what happened, how it was diagnosed, and what would have helped someone else handle it. That creates practical knowledge without pretending that troubleshooting can be completely scripted.
Make the risk visible, document that you’ve raised it, and stop compensating for the staffing decision with your personal time. Cross-training and documentation take real hours, so ask your VP which existing priorities should be delayed to make room for them. Set realistic timelines and explain that training will temporarily reduce productivity before it improves coverage.
For PTO, establish a formal backup plan: a change freeze for risky work, a named primary backup, and the MSP as the escalation path. Unless there is a genuine emergency arrangement in your employment terms, being unavailable while on leave is reasonable. If you keep solving everything from vacation, leadership has no incentive to fix the underlying problem.
That’s essentially what happened this time. I didn’t respond to calls, messages, or emails while away, and management quickly realized how dependent the team was on me. It helped make the problem visible, although it also resulted in a list of new action items being handed back to me.
Cross-training only works if the other people periodically perform the work. A one-hour explanation once a year won’t create real coverage. Consider rotating ownership of infrastructure tasks for a defined period, with you available for escalation while someone else handles the routine work. That gives them experience and reveals which areas need better documentation.
At the same time, don’t accept responsibility for fixing the entire staffing model. Your job is to identify the risk, provide reasonable training and documentation, and clearly state what cannot be covered with the available resources. Management has to decide whether to add staff, fund more MSP support, reduce the workload, or accept the operational risk.
This may also be a warning about the future of the role. When leadership refuses infrastructure headcount but emphasizes using the MSP, they may be considering moving more of the work externally. That doesn’t necessarily mean you’re about to be let go, but it’s worth preparing for.
Keep a professional record of the risks you identified, the coverage gaps, and the work you’ve completed. Update your résumé and quietly explore other options. Don’t rely on being indispensable as job security; sometimes companies respond to a single point of failure by trying to eliminate or outsource the position instead of properly supporting it.

The goal shouldn’t be to make everyone an identical replacement for you. It should be to ensure someone else can safely perform the common tasks, recognize when they’re out of their depth, and escalate with enough information for the MSP or another engineer to help.