AVD Hosts Experiencing Severe Disk Contention During Morning Logins

0
0
Asked By MellowCedar47 On

We have an Azure Virtual Desktop environment with around 163 Windows 11 24H2 Enterprise multisession hosts supporting roughly 2,000 devices. Each host uses an E8s_v5 VM size with 8 vCPUs and 64 GB of RAM, along with a P10 128 GB Premium SSD. The environment was stable until September 23, when the normal morning login period began causing heavy disk activity and severe contention. Users are seeing blank screens for up to 20 seconds before their desktops load, which is affecting the overall experience. Windows updates do not appear to be the cause, and our FSLogix redirection.xml configuration is correct. The disks are still showing as Premium SSD LRS, although Nerdio manages automatic conversion between Standard and Premium storage. Affected VMs have disk queue lengths above 20, while unaffected machines are generally below 2. We have not found one particular process responsible for the disk activity and are also investigating this with Microsoft. Has anyone seen similar behavior or have suggestions for identifying the source?

3 Answers

Answered By NorthstarKite6 On

It may be worth testing a higher disk tier on a small group of hosts. Similar AVD environments have run into Premium SSD IOPS limits during morning logon bursts, sometimes after increasing the frequency of Log Analytics collection. Moving a few machines from P10 to P15, or testing Premium SSD v2 where available, can show whether the problem is storage performance rather than the VM size. Compare the hosts over a full business day before changing the rest.

CopperVale52 -

One investigation also pointed toward Microsoft Defender activity, although the exact trigger was never identified. Upgrading the disk tier removed the contention, so testing the larger SKU first may be a practical way to confirm whether capacity is the limiting factor.

Answered By QuietHarbor8 On

Start by confirming the actual disk SKU, IOPS, throughput limits, and host-level metrics rather than relying only on what the portal reports for the managed disk. A P10 can become the bottleneck during a large burst of simultaneous logins, especially with FSLogix profile activity, Defender scans, and Log Analytics collection happening at the same time. Compare disk queue depth, read/write latency, IOPS, and throughput on affected and unaffected hosts during the login window.

AmberField29 -

The affected machines have queue lengths above 20, while the unaffected ones are below 2. Nerdio has converted the disks between Standard and Premium, but they currently show as Premium SSD LRS, and no single process is obviously dominating disk usage.

Answered By LunarMaple31 On

Check the event logs and performance counters around the first affected login period, including FSLogix, Defender, Windows Search, Windows Update, and the Log Analytics agent. Look for profile-container attach delays, storage warnings, antivirus scan activity, or a sudden increase in telemetry collection. It is also useful to compare a newly provisioned host with an older one to determine whether the change is coming from the image, an agent policy, or the Azure storage platform.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.