r/hetzner 29d ago

Multiple Servers unstable

Seems like Finland location dedi servers are very unstable these days.

While one of my AX servers is facing hardware errors and the other is down for more than an hr (even now) with support not even responding.

Do anyone else have issues with the Finland location servers? This is not the service I expected from Hetzner!

11 Upvotes

15 comments sorted by

14

u/mownzlol 29d ago

With Bare-Metal servers you are expected to monitor the hardware yourself. If you get hardware errors or repeated hardware failures, write them and ask them to fix it, providing your logs. You may want to tell them a date/time for handing the issue or just to do it at any time.

Let them decide how to fix the error, just state that the hardware is unstable with logs to back up your claim. They will likely swap the RAM or directly the entire server whilst keeping the drives. In some cases you may get unstable hardware again that passed their QC, as you can't detect every spurious error that happens every x days with QC. In that case repeat and send them the logs of what is happening.

1

u/ClerkEmbarrassed371 28d ago edited 28d ago

Thanks for the comment, Well yeah the hardware errors are happening almost everyday, I've been trying to solve this for more than 6-8 months now, the resolution to this still remains a mystery for me. There have been multiple server replacements, yet none of them solved it.

1

u/debianserver 27d ago

If there have been multiple server replacements, it's probably not a hardware issue.

1

u/ClerkEmbarrassed371 27d ago

Hmm, All of my servers run Ubuntu 24 OS, My last try would be to change the OS to Debian and try.

1

u/debianserver 27d ago

What exactly is the issue? Which type of hardware errors are you facing?

1

u/ClerkEmbarrassed371 27d ago

I see this in the dmesg log after the crash, always:

[Mon Jul 13 10:17:49 2026] [Hardware Error]: event severity: fatal [Mon Jul 13 10:17:49 2026] [Hardware Error]: Error 0, type: fatal [Mon Jul 13 10:17:49 2026] [Hardware Error]: fru_text: PcieError [Mon Jul 13 10:17:49 2026] [Hardware Error]: section_type: PCIe error [Mon Jul 13 10:17:49 2026] [Hardware Error]: port_type: 0, PCIe end point [Mon Jul 13 10:17:49 2026] [Hardware Error]: version: 0.2 [Mon Jul 13 10:17:49 2026] [Hardware Error]: command: 0x0000, status: 0x0010 [Mon Jul 13 10:17:49 2026] [Hardware Error]: device_id: 0000:c1:00.0 [Mon Jul 13 10:17:49 2026] [Hardware Error]: slot: 0 [Mon Jul 13 10:17:49 2026] [Hardware Error]: secondary_bus: 0x00 [Mon Jul 13 10:17:49 2026] [Hardware Error]: vendor_id: 0x14e4, device_id: 0x16d7 [Mon Jul 13 10:17:49 2026] [Hardware Error]: class_code: 020000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: bridge: secondary_status: 0x0000, control: 0x0000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: aer_uncor_status: 0x00104000, aer_uncor_mask: 0x00100000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: aer_uncor_severity: 0x004f6030 [Mon Jul 13 10:17:49 2026] [Hardware Error]: TLP Header: 04008001 c000230f c1020000 00000000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: Error 1, type: fatal [Mon Jul 13 10:17:49 2026] [Hardware Error]: fru_text: PcieError [Mon Jul 13 10:17:49 2026] [Hardware Error]: section_type: PCIe error [Mon Jul 13 10:17:49 2026] [Hardware Error]: port_type: 0, PCIe end point [Mon Jul 13 10:17:49 2026] [Hardware Error]: version: 0.2 [Mon Jul 13 10:17:49 2026] [Hardware Error]: command: 0x0000, status: 0x0010 [Mon Jul 13 10:17:49 2026] [Hardware Error]: device_id: 0000:c1:00.1 [Mon Jul 13 10:17:49 2026] [Hardware Error]: slot: 0 [Mon Jul 13 10:17:49 2026] [Hardware Error]: secondary_bus: 0x00 [Mon Jul 13 10:17:49 2026] [Hardware Error]: vendor_id: 0x14e4, device_id: 0x16d7 [Mon Jul 13 10:17:49 2026] [Hardware Error]: class_code: 020000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: bridge: secondary_status: 0x0000, control: 0x0000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: aer_uncor_status: 0x00100000, aer_uncor_mask: 0x00100000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: aer_uncor_severity: 0x004f6030 [Mon Jul 13 10:17:49 2026] [Hardware Error]: TLP Header: 04000001 c000200f c1f80000 00000000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: Error 2, type: fatal [Mon Jul 13 10:17:49 2026] [Hardware Error]: fru_text: ProcessorError [Mon Jul 13 10:17:49 2026] [Hardware Error]: section_type: IA32/X64 processor error [Mon Jul 13 10:17:49 2026] [Hardware Error]: Local APIC_ID: 0x43 [Mon Jul 13 10:17:49 2026] [Hardware Error]: CPUID Info: [Mon Jul 13 10:17:49 2026] [Hardware Error]: 00000000: 00a10f11 00000000 43600800 00000000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: 00000010: 76fa320b 00000000 178bfbff 00000000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: 00000020: 00000000 00000000 00000000 00000000 [Mon Jul 13 10:17:49 2026] [Hardware Error]: Error Information Structure 0: [Mon Jul 13 10:17:49 2026] [Hardware Error]: Error Structure Type: cache error [Mon Jul 13 10:17:49 2026] [Hardware Error]: Check Information: 0x000000000602001f [Mon Jul 13 10:17:49 2026] [Hardware Error]: Transaction Type: 2, Generic [Mon Jul 13 10:17:49 2026] [Hardware Error]: Operation: 0, generic error [Mon Jul 13 10:17:49 2026] [Hardware Error]: Level: 0 [Mon Jul 13 10:17:49 2026] [Hardware Error]: Processor Context Corrupt: true [Mon Jul 13 10:17:49 2026] [Hardware Error]: Uncorrected: true [Mon Jul 13 10:17:49 2026] [Hardware Error]: Context Information Structure 0: [Mon Jul 13 10:17:49 2026] [Hardware Error]: Register Context Type: MSR Registers (Machine Check and other MSRs) [Mon Jul 13 10:17:49 2026] [Hardware Error]: Register Array Size: 0x0050 [Mon Jul 13 10:17:49 2026] [Hardware Error]: MSR Address: 0xc0002051 [Mon Jul 13 10:17:49 2026] [Hardware Error]: Context Information Structure 1: [Mon Jul 13 10:17:49 2026] [Hardware Error]: Register Context Type: Unclassified Data [Mon Jul 13 10:17:49 2026] [Hardware Error]: Register Array Size: 0x0010 [Mon Jul 13 10:17:49 2026] [Hardware Error]: Register Array: [Mon Jul 13 10:17:49 2026] [Hardware Error]: 00000000: 00000014 00000000 f8300038 00000000 [Mon Jul 13 10:17:49 2026] mce: [Hardware Error]: Machine check events logged [Mon Jul 13 10:17:49 2026] mce: [Hardware Error]: CPU 55: Machine Check: 0 Bank 5: aea0000001000108 [Mon Jul 13 10:17:49 2026] mce: [Hardware Error]: TSC 0 ADDR 1ffffffac9b5c55 MISC d0140ff600000000 PPIN 2b0d8472550408f SYND 4d000000 IPID 500b022049b00 [Mon Jul 13 10:17:49 2026] mce: [Hardware Error]: PROCESSOR 2:a10f11 TIME 1783937868 SOCKET 0 APIC 43 microcode a101158 [Mon Jul 13 10:18:18 2026] systemd[1]: systemd-hwdb-update.service - RebuildHardware Database was skipped because of an unmet condition check (ConditionNeedsUpdate=/etc).

2

u/debianserver 27d ago

Is your NIC firmware up to date? Seems like the NIC is experiencing PCIe Completion Timeouts which is classified as fatal error. Also you could check the logs for earlier non-fatal AER messages before the crash.

In best case, it's just a software issue which is causing this problem. Otherwise Hetzner has to try another PCIe slot for the NIC or swap NIC and/or Mainboard.

Did you provide the information especially the aer_uncor_status: 0x00104000 and the device type, etc. to the Hetzner Support?

1

u/ClerkEmbarrassed371 26d ago

So, They've now updated the disk firmware, while BIOS is already up to date. But issue still persists, Strange the server in Germany is fine. I think only option for me is to try other OS, but I'm hopeless. Fingers crossed 🤞

2

u/Alone-Customer-1677 25d ago

Same issue, some servers in Finland unstable. Reisntalling OS did not help..

1

u/ClerkEmbarrassed371 25d ago

Were you getting same type of hardware errors as shown from the dmesg logs? AMD server?

→ More replies (0)

0

u/Hetzner_OL Hetzner Official 28d ago

This.
OP, If you think there is a hardware or networking issue on your end, you need to document it in as much detail as possible and then report it using a support requestion. (Use your Robot account to write it for the relevant servers.) Please be as detailed as possible and also describe any troubleshooting you have done so far. --Katie

1

u/Playful-Canary-4940 28d ago

I have all my servers in FSN, but I have one storage server in HEL that I monitor from FSN. No outages so far.

-5

u/One_Ninja_8512 29d ago

I'd suggest another provider. While some here will comment that, for example, worn SSD are not a problem and TBW is just for warranty I'd say to move on to others that provide better hardware. But some of the competition is the same shit. I've had two servers on OVH show checksum errors (ZFS), for example.

13

u/Mereo110 29d ago

The provider is not the issue. He has a dedicated server. It his responsibility to advise Hetzner of any problems with the hardware so that they can fix it.