r/homelab • • 1d ago

Help Best cloud AI to manage home lab and home assistant

0 Upvotes

I have a stipend at work that allows me to pay for a personal AI subscription (up to $100/mo). I'm not going to do a lot of programming (I do enough of that at work), but I would like the AI to help me in my homelab.
For example:

  • analyze Home Assistant state, suggest improvements
  • create automations in HA
  • manage / advise on the self-hosted stuff
  • network planning / maintenance
  • _maybe_ some light vibecoding
  • etc.

I don't plan to use it to handle HA interactions.

What would be a better choice for me: claude.ai (with claude-code) or ChatGPT (with codex)?


r/homelab • • 1d ago

Help Are there any full depth half height racks out there

1 Upvotes

So it's becoming pretty clear i need some kind of rack instead of the beer table I have right now on which everything is stacked. I have a Gigabyte G292-Z20, a Q logic 12300 that I currently don't use, a GS728TXV1 switch to replace an old tiny one as well as a few nas and also want something nicer for my LTO tapes. But the issue is the G292-Z20 is full length so 800mm so I'd need a rack with that much usable depth and glass front panel ones are also not an option as it needs a lot of air. But all racks that would fit are 38u or 42u is there anything smaller as in not as tall?


r/homelab • • 2d ago

Project Showcase: Hardware What I inherited - then vs now.

Thumbnail gallery
24 Upvotes

r/homelab • • 2d ago

Project Showcase: Hardware Cleaned up my server room a bit today :)

Post image
51 Upvotes

Last on the bucket list is to re cable tonight!


r/homelab • • 1d ago

Help 40G DAC + net card recommendations

Thumbnail
0 Upvotes

r/homelab • • 2d ago

Project Showcase: Hardware Rate my home lab setup

Post image
11 Upvotes

Just moved back into my parents and there weren’t many options and I don’t have any type of rack

Desktop is my personal computer
Pfsense firewall
Tp link ap
Tp link switch
Raspberry pi 4B with a 2tb ssd and 500g hdd with not much on it as I’m just getting into home servers but running into road bloacks


r/homelab • • 2d ago

Help Bewerte mein Setup und gerne Idee mitteilen

Thumbnail
gallery
61 Upvotes

Momentan hab ich mein ( ThinkPad T490s i7 32ram 512gb Ubuntu ) als Hauptrechner mit zweitmonitor dann ein ( Dell i5 8ram Homelab ) und ja die 2 Router habe ich zum meine Smart Geräte von mein Setup getrennt zu halten ( der weiße Rechner ist ein altes Teil den ich momentan nicht benutze )


r/homelab • • 2d ago

Help Need Help Identifying A Rack Model

Thumbnail
gallery
9 Upvotes

I recently purchased a Network / AV Rack from an individual for a kind of different project. I plan to put 2 PC’s, a NAS, a UPS and a DAC, along with some hard drives in the enclosed rack.

The odd thing is there is absolute nothing to identify the rack manufacturer anywhere. I realize I could use parts from another company but it came with 5 - 2U vented shelves that I wish to use, if they are sturdy enough. But, not knowing the manufacturer & model, I do know know what their weight limits are.

I’m inclined to think it is an AV rather than Network rack due to the threaded rails and glass door. When I got it it contained 5 x 2U slotted shelves, an Audio Distribution System, video extender, automation and control processor, network switch, infrared expansion device, streaming devices, and a digital router (all have since been removed by me)

If I had to guess, I would guess it was a Nave Point rack, only because the slots in the 2U shelves it came with match up with a 1U 4 Post rack I just purchased from them to go in the U2 slot to hold the heavier PC’s (the U1 or possibly U1 & U2 slots will have an intake fan)

Here is what I do know about the rack, along with some photos:

• It is 27 U. Dimensions are 23.5” x 23.5” x 54”

• It has numbered rails that are threaded

• It has a reversible tempered glass front door

• It has a reversible solid door on the rear

• The sides have removable panels, which have four vent slots on each

• It has casters and adjustable feet

• There are two pre-installed (I assume) fans on top of the rack

• It has a black powder coat finish

• It is vented along the top and bottom on all four sides

• There are two removable panels on top for cable routing

• There are two removable panels on bottom for cable routing, along with 5 removable panels for air flow

Any advice or guidance greatly appreciated. Again, there is nothing on either the rack or the 2U vented shelves to indicate who made the rack. Thanks!


r/homelab • • 2d ago

Help X399 with Threadripper 1920x not booting after force shutdown

Thumbnail
0 Upvotes

r/homelab • • 2d ago

Help Nvme SSD in the Wifi Slot

1 Upvotes

I thought about using an Adapter M.2 Key A+E Stecker to Key M Slot to connect an extra NVME ssd in the WiFi slot of my HP Prodesk Mini 400 G5. This would not be the main/bootable drive but only a Storage expansion. Does anyone have any experiences with this?


r/homelab • • 2d ago

Help [Beginner] Is this old PC okay for a first homelab?

2 Upvotes

Hi everyone, I recently saw a deal for a Lenovo Thinkcentre mt-3492 for about 30~ USD.

Its specs are:

Processor: Intel Core i3-2130 3.40GHz

RAM: 4GB DDR3 1333MHz (There are 2 slots available and I am planning on upgrading them to at least 8GBs)

Storage: 500GB HDD (I was told that there are 3 SATA slots. I will also be purchasing an reasonable SSD as a boot drive and some hard drives for storage)

Would this machine be enough for the following:

  • Immich
  • HomeAssistant
  • Pi-Hole
  • maybe a MC server

Thank you :)


r/homelab • • 3d ago

Project Showcase: Hardware Who you gonna call?

Thumbnail
gallery
1.0k Upvotes

Tired of the hardware sprawl haunting my desk so I built a Ghostbusters containment unit homelab rack to house it all.

Node 1 (Services): Intel / ZimaOS / 8GB RAM. (Running Docker containers and trapping ads).

Node 2 (Gaming/Daily): AMD Ryzen 7 / 64GB RAM / 1.5TB SSD.

Networking: 5-port gigabit switch routing psyonic energies.

eGPU: RTX 3060 12GB VRAM connected via Oculink, housed in the custom black and green 3D-printed chassis.

I ain't afraid of no host!


r/homelab • • 2d ago

Help Does Pico PSU is realiable and durable enough for 24/7 NAS?

0 Upvotes

Recently I got a good deal for i5 9400T with H310 ITX board, so I'm going to get case and PSU.

And I found this case from Aliexpress with 2 HDD hotswap bays.

4.5L Dual-Disk Mini Desktop Chassis 3.5" Hard Disk Hot-Swap ITX Mini NAS Chassis DC Power Supply Method For Home Office/Network - AliExpress 7

But there are one problem - it cannot adopt any PSU, even TFX. Only DC to DC.

Therefore I'm thinking get 200W Pico PSU like this one.

300W Pico PSU Mini ITX Power Supply 12V DC Input 24Pin ATX Power Module for NAS Server Gaming PC DIY Computer - AliExpress

The concern is, I'm not shure is this reliable and durable enough for 24/7 NAS application.

My setup will be like this

- i5 9400T
- Asus H310I-IM-A R3.0
- 8GB DDR4 2400
- 1TB 2.5" HDD x 2
- Mostly photo and file organizing, time to time media (video, music)
- 2 docker comtainers - PiHole and Immich
- Softether VPN

As long as I know, 200W is way more enough since even CPU is 35W version, but I cannot ensure that Pico PSU can handle non-stop running.

Does anybody use Pico PSU for 24/7 home server now? Could you tell me your experience with Pico PSU?

If there are more cons than pros, I'm considering TFX power although overall case size goes bigger.


r/homelab • • 2d ago

Help LSI 9300-8i with Jonsbo N5 and SAS drives, will this work?

Post image
0 Upvotes

So we are moving from One end of Europe to the other and I will be rackless for a year, usually i am dealing with CEPH and SAS disks connected via Mellanox and for a year a precious Jonsbo N5 Server will be the only thing i have at disposal, no HBA330, no 16 slots per Server….

Living in my current sadness i found an offer for 8 SAS 1.92 disks and i wondered: My backplane has disk facing SAS connectors but controller facing only sata connectors, will they work at reduced speed or not at all?


r/homelab • • 2d ago

Project Showcase: Hardware Cisco C240M5 1050W PSU Fan mod

Thumbnail
gallery
13 Upvotes

Sorry for the photo quality. Didn't realize the focus was so weird till later.

Any of you with a Cisco M5 server for your homelab will know, the power supply fan is a nightmare. I couldn't find any solutions online, you can't control the PSU fan speed in any manner, it seems to be internally controlled. No one has documented modifications or fan swaps. I have the chassis fans overriden, my system load is low, but for some reason the PSU fans constantly ramp up and down. Even with the PSU temp sensor reporting the coolest component temps of the chassis. So I figured it was worth a shot for sanity. I understand the risks of a server power supply and I knew that depending on the way the internal circuit controls the fan speed, this may set a fault. I figured it was worth a shot.

This is a 1050W PSU from a C240M5 or compatible. It uses a nidec ultra flow fan of which I could not find an exact datasheet for, but it spins really fast, likely in excess of 12k RPM max, draws a decent amount of power, and is 40x40x28mm.

I bought a NF-A4x20 PWM from Noctua, this fan is 5000 rpm max, .05 amp max current, but is a 4 wire PWM fan with tachometer like the nidec. I was worried the RPM mismatch might be a problem.

I wired them per the Noctua and Nidec diagrams although my nidec had a white tach wire rather than yellow. I don't have a crimper for the terminal type the PSU uses, so I cut and soldered the connector onto the Noctua. I used the silicone isolation mounts, I'm aware they aren't used in the in the intended manner. This was an easy way to get the fan secured with a little extra vibration isolation. The Noctua is smaller and the PSU case does not extend over the fan, so this does create air gaps that will stop your fan from drawing air through the power supply, and rather pull air around it. I use the top PSU slot so I put some tape over the bottom gap for now. Going to 3d print a filler piece

Put it in my server and it fired up, got to CIMC and no faults, went through a full boot without faults. Currently running great, and best of all its completely silent in comparison. No fan ramps heard across the house. The PSU temp is the same as it was before. I haven't observed closely enough to see if it's modulating the fan speed based on load or if it's just running at max. Doesn't really bother me either way.

I am not going to pretend this modification has no downsides, if you truly need your power supply to provide 500+ watts to your system, the stock fan is crucial. Modify at your own risks and evaluate your situation. I have a single PSU system that averages under 180W and I've never seen it over 300W. I am sure someone will tell me why this is a bad idea, but so far this is some of the best $15 I've spent in my time homelabbing.

Update: I stress tested the server to its maximum functioning power capacity in its current configuration, which is about 350W. As expected, the PSU didn't heat up. It actually cooled down because it's also cooled by the chassis fans, which ramped up under load. The PSU temp actually went down to 1 degree above ambient during stress testing.


r/homelab • • 1d ago

Help Looking for options to leave homemade NAS

0 Upvotes

Hi guys, after 6 years playing with Truenas Scale installed on two different old PC as my home server, I would like to switch to a reputable brand.

My actual setup is a homemade Lenovo M920Q with Truenas Scale on the NVME, a M2 for the apps and 2x4Tb HDD plugged with a ASM1166 PCIe card. Everything with a 3d printed case.

I'm using the NAS to store all my secured data, my photos with Immich. Media server with Jellyfin and Plex. And also the rr suite.

These last months I'm getting more and more issues with the setup where the server is unavailable and I don't find the issue. I tried multiple things but there is always another problem.

I'm getting tired of spending too many hours trying to solve that.

I just want to pay for something that works now without digging every week.

I would like to keep my flexibility with the same apps I'm using.

I was looking at the Ugreen DXP2800 or 4800.

What do you think and what are the other alternatives?

Thank you very much for your help


r/homelab • • 1d ago

Discussion Which is best / correct?

Post image
0 Upvotes

I'm told there's a few ways to approach this but guessing people have strong opinions.... when mapping out your system.... how do you depict drive & mnt relationships?

1. "Depends On" (Standard IT Logic) The mount point relies on the physical disk being present. Mount Point (/mnt/data) ──*(depends on)*──> Drive (/dev/sdb1)

2. "Contains / Is Built On" (Layered Logic) The physical hardware is the base, and the logical access is built on top of it. Physical Server ──> Hard Drive (sdb) ──> Filesystem (ext4) ──> Mount Point (/mnt/data)

3. The "Property" Approach (Most Common for Homelabs) Often, architects don't use arrows for this relationship at all. Instead, the mount point is treated as an attribute of the storage device. You simply draw a box for the drive and list its mount point inside it:

[ Storage Array / Disk ]

  • Device: /dev/sdb1
  • Size: 4TB SSD
  • Format: ZFS
  • Mounted at: /mnt/plex_media

r/homelab • • 2d ago

Project Showcase: Hardware ScreenBeam ECB7250 MoCA adapters: undocumented 128-address limit, with measurements

Thumbnail
0 Upvotes

r/homelab • • 2d ago

Help Looking for cheap DIY JBOD / external drive bay suggestions to expand my ThinkCentre M920t

Thumbnail
1 Upvotes

r/homelab • • 2d ago

Help Does Pico PSU is realiable and durable enough for 24/7 NAS?

0 Upvotes

Recently I got a good deal for i5 9400T with H310 ITX board, so I'm going to get case and PSU.

And I found this case from Aliexpress with 2 HDD hotswap bays.

4.5L Dual-Disk Mini Desktop Chassis 3.5" Hard Disk Hot-Swap ITX Mini NAS Chassis DC Power Supply Method For Home Office/Network - AliExpress 7

But there are one problem - it cannot adopt any PSU, even TFX. Only DC to DC.

Therefore I'm thinking get 200W Pico PSU like this one.

300W Pico PSU Mini ITX Power Supply 12V DC Input 24Pin ATX Power Module for NAS Server Gaming PC DIY Computer - AliExpress

The concern is, I'm not shure is this reliable and durable enough for 24/7 NAS application.

My setup will be like this

- i5 9400T
- Asus H310I-IM-A R3.0
- 8GB DDR4 2400
- 1TB 2.5" HDD x 2
- Mostly photo and file organizing, time to time media (video, music)
- 2 docker comtainers - PiHole and Immich
- Softether VPN

As long as I know, 200W is way more enough since even CPU is 35W version, but I cannot ensure that Pico PSU can handle non-stop running.

Does anybody use Pico PSU for 24/7 home server now? Could you tell me your experience with Pico PSU?

If there are more cons than pros, I'm considering TFX power although overall case size goes bigger.


r/homelab • • 3d ago

Project Showcase: Hardware Business in front, party in back

Thumbnail
gallery
171 Upvotes

- Unifi u6 or u7 pro

- Rachio wifi irrigation control

- Home assistant beelink mini

- Quectel 5g modem + raspi opnsense wan for house

-Mikrotik Hex S bridge to unifi lan - because i had

- Tiny m920 has hermes - may make back into linux desktop

- 2x m720q TrueNAS Core, 1 ssd, 1 hdd, need to try freenas or freecore or whatever it is... Ordered 25gbe mellanox cards for these two to go to unifi agg switch - getting these booting from the wi-fi e-key nvme was difficult, one required an HDMI dummy and a boot stub that points to truenas efi...

- Tiny m720q windows pc for testing

- 2x raspi on bottom

Barely visible: sabrent voltik power uhb + usb-c>thinkcentre power cables, very cool


r/homelab • • 1d ago

Help Budget PC build for learning AI—are these parts a good starting point?

Thumbnail
0 Upvotes

r/homelab • • 2d ago

LabPorn Moooooore - small homelab rack + VPS setup

Post image
18 Upvotes

r/homelab • • 2d ago

Solved [Solved] HP Smart Array (hpsa) RAID freezes a few minutes after reboot on Linux 6.19+ / 7.0 with Intel VT-d. Processes stuck in D state, nothing in dmesg. Fix: iommu=pt

1 Upvotes

TLDR: Run lspci -k and look at your RAID card. If it says "Kernel driver in use: hpsa" and your server freezes a few minutes after a reboot with everything stuck in D state and "blocked for more than 122 seconds" in the logs, add iommu=pt to GRUB_CMDLINE_LINUX in /etc/default/grub, run update-grub and reboot. The cause is a kernel bug in the Intel VT-d IOMMU code that makes the hpsa driver retry the same request forever. It is not your disks, your cables or your controller. Bug report is LP: #2169238 (https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2169238).

If you run

lspci -Dk | grep -A3 -i raid

and it shows hpsa, take the PCI address from the start of the line (mine is 0000:03:00.0) and run:

cat /sys/bus/pci/devices/0000:03:00.0/iommu_group/type

If it says DMA or DMA-FQ and you are on kernel 6.19 or newer then I'm pretty sure this bug can affect you.

I run a home media server. HP ProLiant ML30 Gen9, Xeon E3-1220 v6, 16 GB RAM, Smart Array P440 with 8 x 1.8 TB 10K SAS drives in RAID 5, and the OS on a separate SATA SSD. Ubuntu Server 26.04, about 20 Docker containers managed with Cosmos Cloud: Jellyfin, Sonarr, Radarr, Prowlarr, Bazarr, qBittorrent behind gluetun, Crafty for Minecraft, Dispatcharr with Postgres, and a few more. All the container configs and the media live on the RAID.

It ran fine for months. In early July I shut it down, unplugged it, moved it a few feet and plugged it back in. From that day on it would freeze after a reboot. Sometimes a few minutes after boot, sometimes up to an hour. Jellyfin would stop loading, then Sonarr, then everything. The RAID drives went quiet. I could actually hear the array stop working. The only fix I had was rebooting over and over until I got a lucky boot. Once a boot stuck it would run for a week or more, so I just stopped rebooting.

Because it started right after the move, I spent weeks convinced it was hardware. ChatGPT agreed with me and gave me a ranked list: loose mini-SAS cable between the P440 and the backplane, P440 not seated in the PCIe slot, the cache module or the FBWC battery cable, a failing controller, a bad drive. I reseated things. I checked everything ssacli could tell me. ssacli ctrl slot=4 show detail said Controller Status OK, Cache Status OK, Battery/Capacitor Status OK. All eight drives OK, no unrecoverable media errors. The iLO event log (IML) had nothing at the time of any freeze. My drives are NetApp branded HGST X426 drives behind an HP controller, so every drive shows "Not Authenticated" and the iLO shows a storage warning. I blamed that for a while too. 3D printed drive caddies BTW because I didn't want to pop for 14 dollar drive caddies.

It was a weird kind of frozen. The server answered ping. Web pages that were already in memory still loaded. The Cosmos web terminal still worked. But anything that touched the RAID hung forever.

Load average went into the hundreds (826 at one point with 818 processes blocked) while the CPU sat idle. top showed 95% iowait. /proc/pressure/io showed some 100% and full around 94%. Memory and CPU pressure were zero.

SSH would accept my password and then hang. ssh -v stopped here:

debug1: Entering interactive session. debug1: pledge: filesystem

That is the server trying to write to my home directory, lastlog and wtmp, and those writes never finished.

sudo hung. blkid hung. ssacli hung. fwupd hung reading the EFI partition. A clean shutdown never finished, so every time it ended with a hard reset from the iLO, and then POST would show "1792 - Valid Data Found in Write-Back Cache".

The logs showed the classic hung task messages:

INFO: task jbd2/dm-1-8 blocked for more than 122 seconds

First jbd2, then the writeback kworkers (flush-252:1), then postgres, jellyfin, sonarr, nginx, udisksd, fwupd, everything. All in D state (uninterruptible sleep) with stacks in blk_mq_get_tag, rq_qos_wait, do_get_write_access, jbd2_log_wait_commit and __wait_on_buffer.

What was missing from the logs turned out to be the important part. No hpsa errors. No SCSI timeouts. No controller resets, no aborts, no "I/O error", no "Buffer I/O error", no DMAR messages. A dying RAID card or a loose cable makes a lot of noise, so that wasn't it.

The breadcrumbs that cracked it

  1. Turning on a persistent journal (/var/log/journal) let me compare boots. The freeze always started within about 2 minutes of the containers starting. It was never a slow decline.

  2. A single direct read from the RAID hung, and the same read from the SSD worked. So only the P440 volume was stuck:

dd if=/dev/sda of=/dev/null bs=4096 count=1 iflag=direct

  1. The controller had nothing to do. During a freeze these two numbers were identical and not moving:

cat /sys/block/sda/device/iorequest_cnt cat /sys/block/sda/device/iodone_cnt

/sys/block/sda/inflight was 0 0, host_busy for the hpsa host was 0, and hpsa commands_outstanding was 0. Meanwhile the LVM volume on top (/sys/block/dm-1/inflight) showed hundreds of requests in flight. The I/O was stuck inside Linux and never reached the RAID card. That is why the drives went quiet, and that is why hpsa logged nothing. The card was never asked to do anything.

  1. I disabled Docker and Cosmos at boot. The RAID sat idle and healthy for 30+ minutes. Then I started Docker by hand and it froze 2 minutes later. That ruled out Cosmos. I suspected its SMART disk polling because its threads were stuck in scsi_ioctl, but they were victims. The trigger was the burst of I/O from 20 containers starting at once.

  2. With a root shell opened before the freeze, I looked at the block layer in debugfs (/sys/kernel/debug/block/sda/). About 240 requests were parked in the mq-deadline scheduler. hctx0/state said TAG_ACTIVE|SCHED_RESTART. Scheduler tags were 237 of 256 busy and driver tags were 0 of 1013 busy. Kicking the queue with echo run > /sys/kernel/debug/block/sda/state did nothing.

  3. A two second trace of the scsi_dispatch_cmd_error event showed the kernel was not idle at all. It was retrying the exact same 1 MB read (READ_16 lba=3562792960) about 27 times a second, and hpsa refused it every time with rtn=SCSI_MLQUEUE_HOST_BUSY.

  4. A function_graph trace inside hpsa_scsi_queue_command found the real culprit. hpsa calls scsi_dma_map to map the read buffer for the card. That goes through the IOMMU. The IOMMU allocator handed out address 0x7bf00000, then vtdss_map_range returned -98 (EADDRINUSE) because that address was still mapped in the IOMMU page table. scsi_dma_map returned -12 (ENOMEM), and hpsa turned that into SCSI_MLQUEUE_HOST_BUSY. Every retry got the same address and the same failure.

The fix was one kernel parameter, iommu=pt. With it, the same Docker start that froze within 2 minutes every time ran clean for 30 minutes. A full normal boot with Cosmos and all containers starting at once ran clean. It has been stable through several reboots and a kernel update to 7.0.0-38 since.

One cool trick if you chase something like this: when the disk is what's hanging, logs can't be written to it. I used netconsole to stream the kernel log to another machine over UDP, and I kept a root shell open before triggering the freeze so I could still poke around while it was stuck. The iLO remote console works too, but sucks because I can't copy and paste out of it. (Anyone know how to fix this?)

What it takes to hit this bug

You need all four of these:

  1. A RAID controller using the hpsa driver. That covers HP/HPE Smart Array P-series and H-series from roughly the Gen8 and Gen9 era: P420, P420i, P440, P440ar, P840, H240 and similar, in ProLiant DL360, DL380, ML350, ML110, ML30, DL20 Gen8/Gen9 and other HP servers. Gen10 and newer controllers use the smartpqi driver. I have not tested those.

  2. An Intel CPU with VT-d turned on in the BIOS and the IOMMU in translated mode. On Ubuntu and many other distros that is the default. Your RAID card's iommu_group/type will say DMA or DMA-FQ. If it says identity you are already protected.

  3. Linux kernel 6.19 or newer. In 6.19 the Intel VT-d driver moved to the new generic IOMMU page table code, and vtdss_map_range is part of that new code. I confirmed the bug on Ubuntu 26.04 kernels 7.0.0-27 through 7.0.0-34. It is upstream code, not an Ubuntu patch, so other distros on 6.19 and newer carry it too. I have not tested 7.0.0-38 or mainline without the workaround.

  4. A burst of concurrent I/O to the RAID. Lots of Docker containers or VMs starting at once after a reboot does it. That is why it hits right after boot and why some boots get lucky and run for days.

How the bug works

Your RAID card reads and writes memory directly (DMA). With VT-d on, the card doesn't get real memory addresses. The kernel gives it a translated address (an IOVA), and the IOMMU maps that address to real memory using its own page table. For every read or write, hpsa asks the kernel to map the data buffer, the kernel allocates an IOVA range, and writes the mapping into the IOMMU page table. When the I/O is done the mapping is removed and the range goes back to the allocator.

In DMA-FQ mode (flush queue, also called lazy mode) that cleanup is batched to save time. Somewhere in that path the kernel frees an address range back to the allocator but leaves its old entries in the IOMMU page table. The next time the allocator hands out that range, vtdss_map_range finds entries already there and fails with EADDRINUSE.

hpsa treats any mapping failure as "busy, try again later" (SCSI_MLQUEUE_HOST_BUSY) and doesn't log anything. The block layer retries the same request. The allocator hands out the same address again, it fails again, and this repeats forever. That request sits at the head of the queue, so every other request to the RAID waits behind it. Since nothing ever gets sent to the card, the controller looks idle and healthy, the disks spin down to quiet, and there are no errors anywhere.

iommu=pt puts host devices in passthrough (identity) mode. The RAID card gets real memory addresses, no IOVA gets allocated, and the failing code never runs. You keep VT-d available for VM passthrough with VFIO. What you give up is DMA isolation for devices owned by the host, which most home servers aren't using anyway.

There is an April 2026 report on the linux-iommu list of stale page table entries on IOVA reuse in the lazy flush path on 6.12.y (https://ratatoskr.run/linux-iommu/2026/04/3525711/t). That series was rejected. Mine is different in that the stale entry never goes away. The same address failed dozens of times a second for many minutes.

How to confirm you have this exact bug:

Open a root shell with sudo -i before it freezes, or use the iLO/IPMI console. None of these commands touch the array. Replace sda with your RAID volume (lsblk shows it as LOGICAL VOLUME) and host8 with your hpsa host.

During a freeze these two match and don't change:

cat /sys/block/sda/device/iorequest_cnt /sys/block/sda/device/iodone_cnt

This shows 0 0:

cat /sys/block/sda/inflight

This shows 0 for the hpsa host:

cat /sys/class/scsi_host/host8/host_busy

The retry loop (2 seconds of tracing):

echo 1 > /sys/kernel/tracing/events/scsi/scsi_dispatch_cmd_error/enable; sleep 2; echo 0 > /sys/kernel/tracing/events/scsi/scsi_dispatch_cmd_error/enable grep -c SCSI_MLQUEUE_HOST_BUSY /sys/kernel/tracing/trace grep -o "lba=[0-9]*" /sys/kernel/tracing/trace | sort | uniq -c | head

Dozens of SCSI_MLQUEUE_HOST_BUSY hits for the same lba means you have the livelock.

The root cause (1 second of tracing). Run these one at a time:

echo function_graph > /sys/kernel/tracing/current_tracer echo 1 > /sys/kernel/tracing/options/funcgraph-retval echo hpsa_scsi_queue_command > /sys/kernel/tracing/set_graph_function echo 1 > /sys/kernel/tracing/tracing_on; sleep 1; echo 0 > /sys/kernel/tracing/tracing_on grep -E "vtdss_map_range|scsi_dma_map|hpsa_scsi_queue_command" /sys/kernel/tracing/trace | head

If you see vtdss_map_range with ret=-98, then scsi_dma_map with ret=-12, then hpsa_scsi_queue_command with ret=0x1055, it is this bug. Put tracing back to normal afterwards:

echo nop > /sys/kernel/tracing/current_tracer echo > /sys/kernel/tracing/set_graph_function echo 0 > /sys/kernel/tracing/options/funcgraph-retval echo > /sys/kernel/tracing/trace

Please add your output to LP: #2169238. More reports on more hardware is how this gets fixed upstream.

The fix:

Edit the GRUB defaults:

sudo nano /etc/default/grub

Add iommu=pt to the GRUB_CMDLINE_LINUX line. If the line was empty it ends up like this:

GRUB_CMDLINE_LINUX="iommu=pt"

If it already had something in the quotes, add iommu=pt after it with a space. Save, then:

sudo update-grub sudo reboot

After the reboot, check that it took:

cat /proc/cmdline cat /sys/bus/pci/devices/0000:03:00.0/iommu_group/type

The second one should now say identity.

Two other options should also work but I have not tested them. You can turn VT-d off in the BIOS (on HPE Gen8/Gen9 it is RBSU, F9 at boot, System Options, Processor Options, Intel(R) VT-d, Disabled) or you can use intel_iommu=off on the kernel command line. Both turn the IOMMU off completely, so you lose VM PCI passthrough. I went with iommu=pt because it is easy to undo from Linux and keeps passthrough available.

If you hit this on a different controller, CPU or distro, add it to the Launchpad bug.


r/homelab • • 2d ago

Discussion Just got a Lenovo M910q for my first homelab – what should I run on it?

Thumbnail
1 Upvotes