I recently returned from 10 days of PTO and had a one-on-one with my VP, who is also my direct manager. Our team consists of me, a senior infrastructure engineer, and three mid-level security engineers. He said the team needs more cross-training, especially on infrastructure, because things effectively stop when I'm away. There won't be additional infrastructure headcount, so we're expected to rely on internal training and our existing MSP.
I've raised the need for redundancy in my role several times over the past few years, both verbally and in writing, but management never acted on it. I already do some cross-training and am willing to help more, since having others understand the environment is good for everyone. The problem is that I regularly receive texts or calls about work while I'm on vacation, and it feels like the organization has no real coverage plan.
My VP also wants SOPs for nearly everything. I'm happy to document repeatable procedures, workflows, access steps, and known issues, but I don't think it's realistic to write a guide for every possible troubleshooting scenario. Troubleshooting requires experience and judgment, and the time spent creating documentation and training also affects project schedules.
Has anyone dealt with a similar single-point-of-failure situation? What actually helped reduce the workload, improve coverage, and make PTO genuinely uninterrupted?
4 Answers
Stop making yourself the safety net during PTO. Give reasonable notice, leave current status and emergency procedures, and then be unavailable. If the organization has an MSP, that is what it should be used for when the primary engineer is away. If every vacation ends with you fixing problems, leadership has no incentive to fund redundancy.
You don’t need to pretend you’re unreachable if there is a genuine emergency, but routine texts, email questions, and preventable outages should wait until you return. A change freeze during your absence may also be appropriate for systems that nobody else can safely operate.
There’s a balance between refusing all documentation and trying to document your entire thought process. Create useful operational material: architecture diagrams, dependencies, credentials and ownership procedures stored securely, common alerts, rollback steps, known failure patterns, and who to call. After a major incident, hold a short review and capture what happened and how it was resolved.
You can’t teach someone to troubleshoot every novel problem, but you can make sure they know where to look and what actions are safe. Pairing, job rotation, and having another person regularly perform the work will build much stronger coverage than sending coworkers a folder of SOPs.
That’s close to what we already have. There are workflow guides, portal instructions, screenshots, and training sessions. The difficult part is predicting what might fail while I’m away, especially when the issue involves a system the security team rarely touches.
Make the risk and the tradeoffs explicit, then leave the decision with management. Cross-training takes real time, so ask whether they are willing to adjust project timelines and workloads while people learn. The training itself is the workload-reduction plan; it doesn’t happen for free.
Document repeatable procedures and common failure modes, but don’t promise a guide for every possible incident. For troubleshooting, a decision tree, system overview, escalation path, and examples are usually more useful than a giant step-by-step manual. Keep a written record that you raised the staffing and coverage risks.
Exactly. Training someone who only occasionally touches the system is not the same as having a backup who can actually own it. Someone should rotate into the infrastructure role for a meaningful period, with their normal workload adjusted accordingly.
I would also pay attention to the business direction here. Saying no to infrastructure headcount while emphasizing the MSP and demanding that one person make everything transferable can mean leadership is considering outsourcing more of the function. That doesn’t necessarily mean your job is immediately at risk, but it is enough reason to keep your résumé current and understand your options.
Continue doing a professional job, but don’t take personal responsibility for an organizational decision to run lean. If they want more availability, more documentation, or broader ownership, ask for the priorities, time, and compensation that come with it.

I did stay completely offline during my most recent PTO, and that’s when the coverage problem became obvious. They managed to get through it, but afterward the VP turned the issue into a list of tasks for me rather than a staffing plan.