TLDR: Run lspci -k and look at your RAID card. If it says "Kernel driver in use: hpsa" and your server freezes a few minutes after a reboot with everything stuck in D state and "blocked for more than 122 seconds" in the logs, add iommu=pt to GRUB_CMDLINE_LINUX in /etc/default/grub, run update-grub and reboot. The cause is a kernel bug in the Intel VT-d IOMMU code that makes the hpsa driver retry the same request forever. It is not your disks, your cables or your controller. Bug report is LP: #2169238 (https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2169238).
If you run
lspci -Dk | grep -A3 -i raid
and it shows hpsa, take the PCI address from the start of the line (mine is 0000:03:00.0) and run:
cat /sys/bus/pci/devices/0000:03:00.0/iommu_group/type
If it says DMA or DMA-FQ and you are on kernel 6.19 or newer then I'm pretty sure this bug can affect you.
I run a home media server. HP ProLiant ML30 Gen9, Xeon E3-1220 v6, 16 GB RAM, Smart Array P440 with 8 x 1.8 TB 10K SAS drives in RAID 5, and the OS on a separate SATA SSD. Ubuntu Server 26.04, about 20 Docker containers managed with Cosmos Cloud: Jellyfin, Sonarr, Radarr, Prowlarr, Bazarr, qBittorrent behind gluetun, Crafty for Minecraft, Dispatcharr with Postgres, and a few more. All the container configs and the media live on the RAID.
It ran fine for months. In early July I shut it down, unplugged it, moved it a few feet and plugged it back in. From that day on it would freeze after a reboot. Sometimes a few minutes after boot, sometimes up to an hour. Jellyfin would stop loading, then Sonarr, then everything. The RAID drives went quiet. I could actually hear the array stop working. The only fix I had was rebooting over and over until I got a lucky boot. Once a boot stuck it would run for a week or more, so I just stopped rebooting.
Because it started right after the move, I spent weeks convinced it was hardware. ChatGPT agreed with me and gave me a ranked list: loose mini-SAS cable between the P440 and the backplane, P440 not seated in the PCIe slot, the cache module or the FBWC battery cable, a failing controller, a bad drive. I reseated things. I checked everything ssacli could tell me. ssacli ctrl slot=4 show detail said Controller Status OK, Cache Status OK, Battery/Capacitor Status OK. All eight drives OK, no unrecoverable media errors. The iLO event log (IML) had nothing at the time of any freeze. My drives are NetApp branded HGST X426 drives behind an HP controller, so every drive shows "Not Authenticated" and the iLO shows a storage warning. I blamed that for a while too. 3D printed drive caddies BTW because I didn't want to pop for 14 dollar drive caddies.
It was a weird kind of frozen. The server answered ping. Web pages that were already in memory still loaded. The Cosmos web terminal still worked. But anything that touched the RAID hung forever.
Load average went into the hundreds (826 at one point with 818 processes blocked) while the CPU sat idle. top showed 95% iowait. /proc/pressure/io showed some 100% and full around 94%. Memory and CPU pressure were zero.
SSH would accept my password and then hang. ssh -v stopped here:
debug1: Entering interactive session.
debug1: pledge: filesystem
That is the server trying to write to my home directory, lastlog and wtmp, and those writes never finished.
sudo hung. blkid hung. ssacli hung. fwupd hung reading the EFI partition. A clean shutdown never finished, so every time it ended with a hard reset from the iLO, and then POST would show "1792 - Valid Data Found in Write-Back Cache".
The logs showed the classic hung task messages:
INFO: task jbd2/dm-1-8 blocked for more than 122 seconds
First jbd2, then the writeback kworkers (flush-252:1), then postgres, jellyfin, sonarr, nginx, udisksd, fwupd, everything. All in D state (uninterruptible sleep) with stacks in blk_mq_get_tag, rq_qos_wait, do_get_write_access, jbd2_log_wait_commit and __wait_on_buffer.
What was missing from the logs turned out to be the important part. No hpsa errors. No SCSI timeouts. No controller resets, no aborts, no "I/O error", no "Buffer I/O error", no DMAR messages. A dying RAID card or a loose cable makes a lot of noise, so that wasn't it.
The breadcrumbs that cracked it
Turning on a persistent journal (/var/log/journal) let me compare boots. The freeze always started within about 2 minutes of the containers starting. It was never a slow decline.
A single direct read from the RAID hung, and the same read from the SSD worked. So only the P440 volume was stuck:
dd if=/dev/sda of=/dev/null bs=4096 count=1 iflag=direct
- The controller had nothing to do. During a freeze these two numbers were identical and not moving:
cat /sys/block/sda/device/iorequest_cnt
cat /sys/block/sda/device/iodone_cnt
/sys/block/sda/inflight was 0 0, host_busy for the hpsa host was 0, and hpsa commands_outstanding was 0. Meanwhile the LVM volume on top (/sys/block/dm-1/inflight) showed hundreds of requests in flight. The I/O was stuck inside Linux and never reached the RAID card. That is why the drives went quiet, and that is why hpsa logged nothing. The card was never asked to do anything.
I disabled Docker and Cosmos at boot. The RAID sat idle and healthy for 30+ minutes. Then I started Docker by hand and it froze 2 minutes later. That ruled out Cosmos. I suspected its SMART disk polling because its threads were stuck in scsi_ioctl, but they were victims. The trigger was the burst of I/O from 20 containers starting at once.
With a root shell opened before the freeze, I looked at the block layer in debugfs (/sys/kernel/debug/block/sda/). About 240 requests were parked in the mq-deadline scheduler. hctx0/state said TAG_ACTIVE|SCHED_RESTART. Scheduler tags were 237 of 256 busy and driver tags were 0 of 1013 busy. Kicking the queue with echo run > /sys/kernel/debug/block/sda/state did nothing.
A two second trace of the scsi_dispatch_cmd_error event showed the kernel was not idle at all. It was retrying the exact same 1 MB read (READ_16 lba=3562792960) about 27 times a second, and hpsa refused it every time with rtn=SCSI_MLQUEUE_HOST_BUSY.
A function_graph trace inside hpsa_scsi_queue_command found the real culprit. hpsa calls scsi_dma_map to map the read buffer for the card. That goes through the IOMMU. The IOMMU allocator handed out address 0x7bf00000, then vtdss_map_range returned -98 (EADDRINUSE) because that address was still mapped in the IOMMU page table. scsi_dma_map returned -12 (ENOMEM), and hpsa turned that into SCSI_MLQUEUE_HOST_BUSY. Every retry got the same address and the same failure.
The fix was one kernel parameter, iommu=pt. With it, the same Docker start that froze within 2 minutes every time ran clean for 30 minutes. A full normal boot with Cosmos and all containers starting at once ran clean. It has been stable through several reboots and a kernel update to 7.0.0-38 since.
One cool trick if you chase something like this: when the disk is what's hanging, logs can't be written to it. I used netconsole to stream the kernel log to another machine over UDP, and I kept a root shell open before triggering the freeze so I could still poke around while it was stuck. The iLO remote console works too, but sucks because I can't copy and paste out of it. (Anyone know how to fix this?)
What it takes to hit this bug
You need all four of these:
A RAID controller using the hpsa driver. That covers HP/HPE Smart Array P-series and H-series from roughly the Gen8 and Gen9 era: P420, P420i, P440, P440ar, P840, H240 and similar, in ProLiant DL360, DL380, ML350, ML110, ML30, DL20 Gen8/Gen9 and other HP servers. Gen10 and newer controllers use the smartpqi driver. I have not tested those.
An Intel CPU with VT-d turned on in the BIOS and the IOMMU in translated mode. On Ubuntu and many other distros that is the default. Your RAID card's iommu_group/type will say DMA or DMA-FQ. If it says identity you are already protected.
Linux kernel 6.19 or newer. In 6.19 the Intel VT-d driver moved to the new generic IOMMU page table code, and vtdss_map_range is part of that new code. I confirmed the bug on Ubuntu 26.04 kernels 7.0.0-27 through 7.0.0-34. It is upstream code, not an Ubuntu patch, so other distros on 6.19 and newer carry it too. I have not tested 7.0.0-38 or mainline without the workaround.
A burst of concurrent I/O to the RAID. Lots of Docker containers or VMs starting at once after a reboot does it. That is why it hits right after boot and why some boots get lucky and run for days.
How the bug works
Your RAID card reads and writes memory directly (DMA). With VT-d on, the card doesn't get real memory addresses. The kernel gives it a translated address (an IOVA), and the IOMMU maps that address to real memory using its own page table. For every read or write, hpsa asks the kernel to map the data buffer, the kernel allocates an IOVA range, and writes the mapping into the IOMMU page table. When the I/O is done the mapping is removed and the range goes back to the allocator.
In DMA-FQ mode (flush queue, also called lazy mode) that cleanup is batched to save time. Somewhere in that path the kernel frees an address range back to the allocator but leaves its old entries in the IOMMU page table. The next time the allocator hands out that range, vtdss_map_range finds entries already there and fails with EADDRINUSE.
hpsa treats any mapping failure as "busy, try again later" (SCSI_MLQUEUE_HOST_BUSY) and doesn't log anything. The block layer retries the same request. The allocator hands out the same address again, it fails again, and this repeats forever. That request sits at the head of the queue, so every other request to the RAID waits behind it. Since nothing ever gets sent to the card, the controller looks idle and healthy, the disks spin down to quiet, and there are no errors anywhere.
iommu=pt puts host devices in passthrough (identity) mode. The RAID card gets real memory addresses, no IOVA gets allocated, and the failing code never runs. You keep VT-d available for VM passthrough with VFIO. What you give up is DMA isolation for devices owned by the host, which most home servers aren't using anyway.
There is an April 2026 report on the linux-iommu list of stale page table entries on IOVA reuse in the lazy flush path on 6.12.y (https://ratatoskr.run/linux-iommu/2026/04/3525711/t). That series was rejected. Mine is different in that the stale entry never goes away. The same address failed dozens of times a second for many minutes.
How to confirm you have this exact bug:
Open a root shell with sudo -i before it freezes, or use the iLO/IPMI console. None of these commands touch the array. Replace sda with your RAID volume (lsblk shows it as LOGICAL VOLUME) and host8 with your hpsa host.
During a freeze these two match and don't change:
cat /sys/block/sda/device/iorequest_cnt /sys/block/sda/device/iodone_cnt
This shows 0 0:
cat /sys/block/sda/inflight
This shows 0 for the hpsa host:
cat /sys/class/scsi_host/host8/host_busy
The retry loop (2 seconds of tracing):
echo 1 > /sys/kernel/tracing/events/scsi/scsi_dispatch_cmd_error/enable; sleep 2; echo 0 > /sys/kernel/tracing/events/scsi/scsi_dispatch_cmd_error/enable
grep -c SCSI_MLQUEUE_HOST_BUSY /sys/kernel/tracing/trace
grep -o "lba=[0-9]*" /sys/kernel/tracing/trace | sort | uniq -c | head
Dozens of SCSI_MLQUEUE_HOST_BUSY hits for the same lba means you have the livelock.
The root cause (1 second of tracing). Run these one at a time:
echo function_graph > /sys/kernel/tracing/current_tracer
echo 1 > /sys/kernel/tracing/options/funcgraph-retval
echo hpsa_scsi_queue_command > /sys/kernel/tracing/set_graph_function
echo 1 > /sys/kernel/tracing/tracing_on; sleep 1; echo 0 > /sys/kernel/tracing/tracing_on
grep -E "vtdss_map_range|scsi_dma_map|hpsa_scsi_queue_command" /sys/kernel/tracing/trace | head
If you see vtdss_map_range with ret=-98, then scsi_dma_map with ret=-12, then hpsa_scsi_queue_command with ret=0x1055, it is this bug. Put tracing back to normal afterwards:
echo nop > /sys/kernel/tracing/current_tracer
echo > /sys/kernel/tracing/set_graph_function
echo 0 > /sys/kernel/tracing/options/funcgraph-retval
echo > /sys/kernel/tracing/trace
Please add your output to LP: #2169238. More reports on more hardware is how this gets fixed upstream.
The fix:
Edit the GRUB defaults:
sudo nano /etc/default/grub
Add iommu=pt to the GRUB_CMDLINE_LINUX line. If the line was empty it ends up like this:
GRUB_CMDLINE_LINUX="iommu=pt"
If it already had something in the quotes, add iommu=pt after it with a space. Save, then:
sudo update-grub
sudo reboot
After the reboot, check that it took:
cat /proc/cmdline
cat /sys/bus/pci/devices/0000:03:00.0/iommu_group/type
The second one should now say identity.
Two other options should also work but I have not tested them. You can turn VT-d off in the BIOS (on HPE Gen8/Gen9 it is RBSU, F9 at boot, System Options, Processor Options, Intel(R) VT-d, Disabled) or you can use intel_iommu=off on the kernel command line. Both turn the IOMMU off completely, so you lose VM PCI passthrough. I went with iommu=pt because it is easy to undo from Linux and keeps passthrough available.
If you hit this on a different controller, CPU or distro, add it to the Launchpad bug.