Why Are Azure Local Cluster Updates Taking Several Hours?

0
0
Asked By MellowPine47 On

We run three on-premises Azure Local (HCI) deployments, and both Azure Local and SBE updates take much longer than expected. Our three-node HPE DL380 Gen11 cluster, with dual Intel Xeon Platinum 8460Y+ CPUs and 1.5 TB of RAM per host, has completed updates in anywhere from 4 hours 6 minutes to 15 hours 28 minutes. Two-node clusters with dual Xeon Silver 4509Y CPUs and 256 GB of RAM per host typically take 4–7 hours, with one update taking over 10 hours.

Health checks have also occasionally failed. We have found that draining roles from each host, rebooting the hosts one at a time, and then starting the update makes the process more reliable. Even so, we now have to start the update, wait for it to pass the health check, and then check the result the next morning.

The clusters are running Azure Local version 12.26.09.1003.7 on Windows Server 2025 version 24H2, build 26100.33438. Updates are started through PowerShell on one of the nodes. Is this amount of downtime normal, and what diagnostics should we use to identify the slowest part of the process?

3 Answers

Answered By ClockworkMango8 On

The first thing to determine is which update stage is consuming the time. You can inspect the update run with `Get-SolutionUpdate -Id | Get-SolutionUpdateRun`, then review `$run.Progress | ConvertTo-Json -Depth 8`. That should show each step's start and finish time instead of treating the entire update as one operation.

It is also worth running `Get-StorageJob` after each node returns to see how long storage repair takes before the next node is drained. With 1.5 TB of memory, measure a cold boot as well—memory training during POST can add significant time to every reboot.

Answered By QuietFalcon31 On

Your preparation steps are sensible: drain roles cleanly, reboot one host at a time, and confirm the health check before proceeding. I would compare the timestamps for reboot, storage resynchronization, health checks, and update installation. That will tell you whether the delay is caused by firmware or memory-training time, storage repair, validation, or the update package itself. The hardware specifications alone do not explain the full variation between four-hour and fifteen-hour runs.

Answered By SilverMaple_62 On

Unfortunately, multi-hour updates are fairly common with Azure Local. Single-node deployments can take four to five hours per update, and larger clusters may take longer because every node has to be drained, rebooted, and brought back through storage repair and health validation. The unusually long runs are still worth investigating with the progress and storage-job timing rather than assuming the whole process is inherently slow.

BrightHarbor5 -

Some deployments do complete much faster, but that usually depends on the hardware, update path, and which phase is waiting. Getting support involved may be necessary if the same stage repeatedly stalls across several updates.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.