A VM was migrated to another host without advance notice, and the migration paused the underlying hardware—including the accelerated NIC—for about a minute. That brief interruption triggered a cascading failure in a replicated database system and has been extremely difficult to recover from.
The support case was opened as Severity B because the service was impaired but not completely down. The stated initial-response SLA for a Standard support plan is less than four hours. Nearly 25 hours later, the only response has been an automated AI-generated message containing generic troubleshooting suggestions that did not help. No one has called, despite 24/7 phone availability being selected on the case.
Does an AI-generated message count as fulfilling the initial-response SLA, or should the SLA require a response from a human support engineer? Has anyone dealt with a similar escalation or found a reliable way to get a live person involved?
3 Answers
If you’re spending $40,000–$70,000 a month, Standard support may not be enough for the escalation path you need. A higher-tier plan, such as Professional Direct, costs more but provides 24/7 pooled escalation support. Unified support goes further with dedicated contacts and more proactive assistance. That still doesn’t excuse missing the stated response target, but the support tier can make a major difference in getting urgent cases attention.
I had a Severity B production case with the same kind of initial-response target, and it took roughly six days to receive a reply. The eventual advice was essentially to keep retrying and hope the issue cleared up. In practice, an automated troubleshooting message may be recorded as an initial response, even though it’s not useful as an investigation or meaningful human contact. I’d document the missed response and escalate through your account or support-plan contacts.
The support failure is frustrating, but it may be worth reviewing the architecture separately from the provider decision. If a one-minute VM interruption can cause a cascading database failure and a complicated recovery, there may be a resiliency or high-availability gap. Moving providers might improve the support relationship, but it won’t by itself eliminate that single point of failure.

That’s fair, and we are reviewing the resilience design. The immediate issue is that we couldn’t get timely, knowledgeable support while dealing with the outage, which is what pushed the provider decision over the edge.