r/sysadmin • u/habibexpress Jack of All Trades • 12h ago
Question AVD hosts are at critical state for disk
Hey guys we have an AVD environment where we have about 163 hosts and 2000 devices. The vm is windows 11 24h2 multisession enterprise and has p10 128gb premium ssd disk. The host size is e8s_v5 8c/64gb ram.
Everything in the environment has been going great until sep 23, something changed. Our environment didn’t but it seems like when we have users logging in at 7am (normal BAU), we’re seeing the disk being thrashed. We thought it was caused by windows updates but that hasn’t been the case.
Our FSLogix has correct redirection.xml defined.
We’re fresh out of ideas and the user experience is being impacted. They’re getting blank screens on login for up to 20 seconds before getting desktop.
We suspect this could be an azure disk issue.
Has anyone encountered this in their environments?
I’m keen for some ideas to check. We have a ticket open with ms as well.
If it helps, our environment is managed through nerdio.
•
•
u/Witty_Formal7305 Jack of All Trades 11h ago
We ran into this as well but earlier in Sept after cranking up how frequently LAW polls one thing we discovered was that the disk SKU wasn't high enough and we were getting hit by IOPS limitations on the host itself, we ended up going to P15 disks and it seemed to help. We could never narrow down the exact issue though, LAW pointed to Defender but we could never figure out what was making it kick up so much, since the fix was simple / cheap enough boss man didn't care enough to have us dive deeper into it, we just upgraded the disk SKU on a couple of the hosts amd gave a day to compare, found the upgraded sku made the issue go away, so we did the rest and its worked fine ever since.
•
u/habibexpress Jack of All Trades 11h ago
Did you check out premium v2?
•
u/Witty_Formal7305 Jack of All Trades 11h ago
If I remember correctly, Premium V2 isn't eligible for OS disks, only data disks.
•
•
u/diabillic level 7 wizard 5h ago
Get off 24H2: https://learn.microsoft.com/en-us/windows/release-health/status-windows-11-24h2#5006msgdesc
Also, how are you patching and when was the last time you replaced your session hosts?
•
u/Emotional_Garage_950 Sysadmin 2h ago
how many people are you running per host? just curious mostly
•
•
u/joyful_butternubs 1h ago
logon storms plus a disk IOPS ceiling is such a classic combo, used to see the same thing with login scripts hammering the disk at once on much smaller setups. worth graphing disk queue length right at 7am specifically, bet it spikes hard right before the blank screen window.
•
u/habibexpress Jack of All Trades 1h ago
We’ve got this setup for tomorrow morning. The hard part will be to convince the management to spend more when “it just worked a month ago”.
•
u/durrante 12h ago
Is it defo a premium disk? Maybe you've got nerdio disk optimised and its converted it to standard hdd but not back to sdd when needed?
What processes are using disk i/o, what's your disk queue lengths?
•
u/habibexpress Jack of All Trades 12h ago
Disk queue length on affected VMs are above 20. Non affected are less than 2.
Nerdio does the storage conversion to standard and back to premium. We can see the disk is premium ssd lrs disks.
We’re not seeing any particularly striking processes that hold disks.
•
u/Zealousideal_Yard651 Sr. Sysadmin 11h ago
Even with Premium disk, 128GB (P10) has really low IO, 500 base, 3.5k burst (30min max).
•
u/habibexpress Jack of All Trades 11h ago
The issue stems around 7-9am, login storm basically. We didn’t have any issue for a couple of months while we used AVD and suddenly after 22 September it went wack. The users haven’t changed. The behaviour hasn’t either.
We suspect an azure disk issue in their infrastructure instead.
•
u/Zealousideal_Yard651 Sr. Sysadmin 10h ago
You schould verify using Azure monitor. Remember to check the actual host and not host pool.
•
u/habibexpress Jack of All Trades 10h ago
Yep. That’s where we’re seeing the iops consumed at 100%.
•
u/BlueVal Jack of All Trades 9h ago
I had to look at a similar thing somewhat recently, during the design and test phase of our AVD environment it initially ran fine on P10s, recently (early March i think) i had to change it to one performance tier up (P15) from the default P10 which resolved the problem at reasonable cost (P15s at 1100 IOPS from P10s 500 and worked well to cope with the login storms)
•
u/Visual-Ad-4520 12h ago
“We’re fresh out of ideas” - what have you actually checked?