We run an Azure Virtual Desktop environment with roughly 163 Windows 11 24H2 Enterprise multisession hosts supporting about 2,000 devices. Each host uses an E8s_v5 VM size with 8 vCPUs and 64 GB of RAM, along with a 128 GB Premium SSD LRS disk. Since September 23, the environment has experienced heavy disk activity when users begin logging in around 7 a.m. This is affecting the user experience, with some users seeing blank screens for up to 20 seconds before the desktop appears.
We initially suspected Windows updates, but that doesn't appear to be the cause. FSLogix is configured with the expected redirection.xml settings, and the disks currently show as Premium SSDs. Our environment is managed through Nerdio, which can convert disks between storage tiers.
On affected VMs, disk queue length is above 20, while unaffected VMs are generally below 2. We haven't found one specific process consuming an unusual amount of disk I/O. We also have a support case open with Microsoft, but we're looking for additional troubleshooting ideas, particularly around Azure disk performance, IOPS limits, FSLogix, Defender, monitoring, or Nerdio automation.
3 Answers
Since Nerdio changes the disk tier automatically, verify the complete history of those changes rather than only checking the current SKU. Confirm that disks were actually returned to Premium SSD after scaling activity and that the requested IOPS and throughput are what you expect. Also compare Azure Monitor metrics for affected and unaffected hosts during the morning login surge. If the disk is Premium but queue length is still above 20, check whether the VM size or host-level limits are being reached as well.
The Premium SSD SKU may simply be too small for the burst of concurrent logins. One team saw similar behavior after increasing the frequency of Log Analytics polling; Defender and monitoring activity appeared to contribute, but the exact trigger wasn’t identified. Moving the affected hosts to a higher disk tier, such as P15, removed the contention, so it may be worth testing that on a few hosts and comparing queue length, IOPS, throughput, and login times before upgrading everything.
Start by checking the Windows event logs around the affected login window, especially storage, disk, NTFS, FSLogix, profile service, and Defender-related events. Correlate those timestamps with the Azure disk metrics and sign-in failures. A high queue length with no obvious top process can still indicate that the disk is hitting its provisioned IOPS or throughput ceiling rather than a single application causing the load.

It may also be worth comparing Premium SSD v2, since it allows more flexible tuning of performance independently from disk capacity. Test it on a small group first and verify that your management tooling supports the conversion reliably.