What’s the best way to start 100+ Azure VMs efficiently?

0
1
Asked By MellowCedar47 On

I'm measuring how long it takes to start about 100 Azure VMs with the Azure CLI. In my initial tests, starting them in batches of 20 took roughly 30 minutes overall, while batches of 50 took about 46 minutes. These are total wall-clock times for all 100 VMs, not the duration of an individual batch. Increasing concurrency made the complete operation slower, so I'm trying to determine whether Azure Resource Manager or Compute throttling, backend queuing, or another platform limit is involved. My first script also did not use the --no-wait option, so I plan to rerun the tests asynchronously and poll the VM power states until every machine is running. I'd like to hear about recommended concurrency levels, experience starting 100 or more VMs at once, and whether power operations can be throttled or queued without obvious HTTP 429 responses. The eventual environment may contain around 250 VMs.

2 Answers

Answered By BrightKite88 On

First, check whether the start command uses --no-wait. Without it, each CLI call waits for the operation to finish, which can make batch behavior look much worse than the platform’s actual ability to process requests. Submit the start operations asynchronously, then poll the VM states separately. Regional capacity and the specific VM sizes can also affect the results.

MellowCedar47 -

Good catch—the test script was missing --no-wait. I’m going to rerun the measurements with asynchronous requests and compare several concurrency levels before drawing conclusions.

Answered By QuietMaple6 On

The large difference between batches of 20 and 50 could indicate control-plane throttling or backend queuing. Azure may not always surface this as an obvious 429; requests can simply take longer as more operations are submitted. There doesn’t seem to be a single universally reliable concurrency value, so I’d benchmark a range such as 10, 20, 50, and 100 while recording request latency, response headers, provisioning state, and the time each VM actually becomes running. A concurrency level around 15–20 may be a useful starting point, but the best value will depend on the region, subscription, VM sizes, and available capacity.

MellowCedar47 -

That’s what I’m trying to establish. I’ll test different concurrency levels and an all-at-once asynchronous run, then poll until all 100 VMs are ready. Since the final environment may grow to roughly 250 machines, identifying whether the bottleneck is client-side waiting, ARM throttling, or platform capacity is important.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.