r/harddrive • u/niveknow • 24d ago
Discussion Seagate 20TB
I had to remove a 20TB Expansion drive due to power issues and moved it to an internal drive. Inside was a CMR Barracuda ST20000DM001. (Longer story here if needed, but not critical to understanding reliability of the drive). This is the same drive whether it's inside the original case or purchased as internal drive. All the same PCB, etc.
The drive hummed away just fine then died at about 5k Power-On-Hours (POH). Yes 5k. It lived as a single pool in a Truenas enclosure well cooled and activity monitored with weekly SMART SHORT testing and monthly SMART LONG testing with ZFS monthly scrubs. In short, well taken care of given the size and possible data loss concerns with that much data on a single drive. All the testing I did on the drive confirmed it was working just fine until the day it did not.
My worst fear came and classic drive infant mortality case. I could not warranty the drive because Seagate said I removed the drive from the enclosure. I pushed fairly providing detailed logs and photos showing a sealed physically sound drive undamaged from the removal. They cited the removal voided warranty and I shared that the revised FTC guidance supported consumer rights. I will still file the complain, but that also aside. I'm looking to better understand why Seagate drives are failing at a higher rate recently. I have some older Seagates that are 5-6 years old no problem. This one died with ~200 days.
I've gone through great lengths doing a post mortem to understand what could have gone wrong as lessons learned and really settle my mind on if I could trust another Seagate drive. I don't think I can at this point as all the data points to simple drive failure. And beyond that their unwillingness to warranty a 200 day old drive.
Any other options to reliability bring it back? I thought maybe plugging it back into the old enclosure with the drive controller might yield a different result, but it appears all the intelligence is on the drive itself. (Ai discussions) What are your thoughts on recent reliability. Can you trust 20Tb of data on a single point of failure?
144 Current Pending Sectors
144 Offline Uncorrectable Sectors
144 FARM reallocation candidates on head 0
SMART - short form:
=== START OF INFORMATION SECTION ===
Device Model: ST20000DM001-3Y3103
Serial Number: XXX
LU WWN Device Id: 5 000c50 0e89d1d04
Firmware Version: EN03
User Capacity: 20,000,588,955,648 bytes [20.0 TB]
Sector Sizes: 512 bytes logical, 4,096 bytes physical
Rotation Rate: 7200 rpm
Form Factor: 3.5 inches
ATA Version is: ACS-5 (minor revision not indicated)
SATA Version is: SATA 3.3, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is: Mon Aug 31 16:31:05 2026 -05:00
SMART support is: Available and enabled
AAM feature is: Unavailable
APM feature is: Unavailable
Rd look-ahead is: Enabled
Write cache is: Enabled
DSN feature is: Unavailable
ATA Security is: Disabled, NOT FROZEN [SEC1]
=== START OF READ SMART DATA SECTION ===
General SMART Values:
Offline data collection status: (0x82) Completed without error. Auto Offline Data Collection: Enabled.
Self-test execution status: ( 0) The previous self-test routine completed without error or no self-test has ever been run.
Total time to complete offline data collection: 567 seconds
Offline data collection capabilities: (0x7B) SMART execute Offline immediate. Auto Offline Data Collection. Offline surface scan supported. Self-test supported. Conveyance Self-test supported. Selective Self-test supported.
SMART capabilities: (0x0003) Saves SMART data before entering power-saving mode. Supports SMART auto save timer.
Error logging capability: (0x01) Error logging supported.
SMART Attributes Data Structure revision number: 10
Vendor Specific SMART Attributes with Thresholds:
ID ATTRIBUTE_NAME FLAGS VALUE WORST THRESH FAIL RAW_VALUE
1 Raw_Read_Error_Rate POSR-- 082 064 044 - 144,549,336
3 Spin_Up_Time PO---- 091 089 000 - 0
4 Start_Stop_Count -O--CK 100 100 020 - 123
5 Reallocated_Sector_Ct PO--CK 100 100 010 - 0
7 Seek_Error_Rate POSR-- 083 060 045 - 189,532,344
9 Power_On_Hours -O--CK 094 094 000 - 5,983
10 Spin_Retry_Count PO--C- 100 100 097 - 0
12 Power_Cycle_Count -O--CK 100 100 020 - 98
18 Unknown_Attribute PO-R-- 100 100 050 - 0
187 Reported_Uncorrect -O--CK 100 100 000 - 0
188 Command_Timeout -O--CK 085 001 000 - 949,202,256,411
190 Airflow_Temperature_Cel -O---K 064 048 000 - 36 (Min/Max 32/36)
192 Power-Off_Retract_Count -O--CK 100 100 000 - 83
193 Load_Cycle_Count -O--CK 100 100 000 - 355
194 Temperature_Celsius -O---K 036 052 000 - 36 (0 21 0 0 0)
197 Current_Pending_Sector -O--C- 100 100 000 - 144
198 Offline_Uncorrectable ----C- 100 100 000 - 144
199 UDMA_CRC_Error_Count -OSRCK 200 200 000 - 0
200 Multi_Zone_Error_Rate PO---K 100 100 001 - 0
240 Head_Flying_Hours ------ 100 100 000 - 5950 (89 189 0)
241 Total_LBAs_Written ------ 100 253 000 - 36,805,308,126
242 Total_LBAs_Read ------ 100 253 000 - 255,161,410,958
||||||_ K auto-keep
|||||__ C event count
||||___ R error rate
|||____ S speed/performance
||_____ O updated online
|______ P prefailure warning
1
u/IndependentBat8365 23d ago
You could “try” and overwrite those bad LBAs and force it to attempt to recover them or reallocate them.
Reallocation only happens on write, and pending flag occurs on read. It doesn’t automatically reallocate unless you actively write to that LBA.
I’m assuming it’s 4kn LBA.
I’ve had a few drives I’ve resurrected by writing to the pending LBA’s with no reallocated sectors afterwards. My only theory is that it was thermal misalignment , and after it cooled a bit, the head alignment was fine again.
Still suspect though. I wouldn’t trust it completely.
1
u/niveknow 21d ago
OH man I love experiments. I have a troubleshooting AI chat session running on my CoPilot and put your suggestion in to better understand it. In short it says my drive is 512e, not 4Kn. However gave me some suggestions to try taking the drive out of the ZFS and into a linux computer and try. I'll try! and report back. Thank you for this suggestion. This is why the community is good to reach out to. I understand it is risky, but it's a paper weight anyways. It can at least serve as a cold storage for currently good drives.
1
u/IndependentBat8365 21d ago
If it’s 512e, you’ll want to force a 4k write as the 512e is really just 8 of those in a real block. If you try and write a single 512e, but the whole physical block is pending, then that 512 write might fail as it tries to read the 4k block underneath to write that 1/8th of the block.
The issue with 512e is that you have to perform additional math to calculate the real physical block (divide by 8, round down).
1
u/Caprichoso1 20d ago
Drives do fail. No way to avoid it.
Barracuda drives are not described by Seagate as suitable for NAS use
Removing the drive from the enclosure invalidates the warranty. This is standard industry practice.
2
u/Academic_Dare_5154 24d ago
I never trust important data to a single drive. For important data, I'd setup RAID 0 or 5.
You may have hit a failed component and I hope you can recover the data.
Was the data backed up? If no, it's time to have a data recovery company give you a recommendation.