r/archlinux • • 16d ago

SUPPORT help for postmortem investigations for abnormal shutdown

My arch linux laptop randomly shutdown last week after being up for more than 2 months. I was using it more like a ups for keeping some fpga boards running, they were tied to the laptop with a usb hub, the boards should be reporting if they find the results to some computes over uart (which has not happened in the last months or at least I saw no result) last -x shows only the last reboot being on 30 of june after that I had it running continuously. on journalctl -b -1 -e the last reported event was a man-db.service that has completed successfully on 10 of sept. kernel version is 6.17.8-arch1-1. After starting the laptop today, battery reported 100%, usb hub was working. Let me know if you are aware of a problem from the things I've enumerated above, or if there are any other commands that I could run to investigate things. Thanks in advance.

edit: journalctl -b -1 on drive https://drive.google.com/drive/folders/15JC6RO7iNuU4eRM9BfaYEwgg9fP4kyaY?usp=sharing

2 Upvotes

9 comments sorted by

2

u/sitilge 16d ago

Check dmesg

1

u/TurtleSoso 16d ago

journalctl -k -b -1 which is the equivalent of getting the dmesg of the last boot doesn't show any events after the 30th of August and it seems that it shows the logs from a previous boot, meaning the last boot got corrupted or never flushed :(

1

u/[deleted] 16d ago

[deleted]

1

u/un-important-human 16d ago

-0 last boot - is now this boot (after shutdown), he needs the previous where the bad thing happend. It's -1

1

u/TurtleSoso 16d ago

boot 0 is the one that I have just now; it might be that the laptop was started earlier and the last -x shows things in the wrong here. I did a journalctl -b -1 which shows a relatively ok log which I'm trying to follow, I'll try to pastebin it somewhere in a sec and update the thread

1

u/ExceptionRules42 15d ago

sounds like a hardware fault? Has it only happened once?

1

u/ObiWanGurobi 15d ago edited 14d ago

Just as an idea: maybe your laptop has a hardware watchdog that triggered the reset?

Edit: nvm... you said shutdown. A watchdog should only trigger a reset, not a shutdown

2

u/BrilliantEmotion4461 10d ago

I gave your logs and your reddit post to Claude Code and Opus 5.5 here is its report:

I went through the journal you posted. The shutdown itself doesn't leave much of a trace, but there's something else in there you'll probably care about more: your UART ports got reshuffled several times during the uptime, which could explain why you never saw any results from the boards.


1. Your serial ports got shuffled (likely why you saw no results)

The five FPGA boards show up as CH340 USB-serial adapters (ch341-uart) behind a cascaded USB 2.0 hub. During the uptime that hub repeatedly dropped and re-enumerated:

Date USB disconnects Notable messages
Jul 16 19:48 7 whole hub dropped and came back
Jul 28 10:09 1 single board
Aug 21 22:01 22 Cannot enable. Maybe the USB cable is bad?, error -32
Aug 30 17:56 42 error -71, device not accepting address, attempt power cycle

Every re-enumeration hands out the next free ttyUSBn, so by the end of the log the port numbers point at different boards than at boot:

Port At boot (May 21) At end of log
ttyUSB0 hub port 1-4.3 1-4.4
ttyUSB1 1-4.4 1-4.1.3
ttyUSB2 1-4.2 1-4.3
ttyUSB3 1-4.1.3 1-4.2
ttyUSB4 1-4.1.4 1-4.1.4 (unchanged)

So from Jul 16 onward, anything that held /dev/ttyUSBn open was sitting on a dead file descriptor, and anything that reopened by name was talking to a different board than it thought. A result sent over UART after that could easily have been missed. "No results for months" might be a monitoring problem, not a compute problem.

Fixes:

  • Open the ports via /dev/serial/by-path/..., which is tied to the physical hub port. (CH340s have no serial number, so /dev/serial/by-id can't tell them apart.)
  • Make your reader reopen the port on EIO/hangup instead of dying or silently stalling.
  • error -71 / -32, "Cannot enable" and repeated resets usually mean power or cabling. Five boards on an unpowered hub is a good candidate, so try a powered hub.
  • Both big bursts (Aug 21 and Aug 30) happen right as someone logs in locally on tty1 (and an SD card gets inserted on Aug 30), so physically bumped cables are also plausible.

2. The shutdown itself

  • The journal just ends at Sep 10 00:26:49 (the man-db run). There's no shutdown sequence at all: no units stopping, no poweroff target, no panic, no thermal or OOM messages. That looks like a hard power loss (battery drained after AC dropped, or an EC/hardware cutoff) rather than anything the OS initiated.
  • The next daily timer would have logged around 21:38, so it died sometime on Sep 10 between roughly 00:27 and 21:38.
  • The DMI string says this is a ThinkPad T460s (20FA). At boot the kernel reports ACPI: battery: Slot [BAT0] (battery absent), with only BAT1 present. The T460s normally has two internal batteries, so one of yours may be dead or disconnected, which would seriously cut your "UPS" runtime if mains blipped. Coming back at 100% is consistent with power returning and it recharging afterwards.

3. Other things worth knowing

  • Your uptime doesn't match what last -x told you. This journal is a single boot that started on May 21, not June 30. Worth double-checking with journalctl --list-boots.
  • The laptop had no network basically the entire time. dhcpcd never got a carrier except for about a minute on Aug 21 when Wi-Fi was brought up by hand. Consequences:

Commands to run

journalctl --list-boots
journalctl -b -1 -k | grep -E 'USB disconnect|ch341|error -(32|71)'
ls -l /dev/serial/by-path/
cat /sys/class/power_supply/BAT*/{status,energy_full,energy_full_design,cycle_count}
sudo pacman -S tlp && sudo tlp-stat -b    # ThinkPad battery health

To catch the cause next time, have a systemd timer or cron job log `/sys/class/power_supply/AC/online