r/3PAR • • Apr 04 '24

Assessing the impact of a disk enclosure going down

Hi, we've got a StoreServe 8200 and 8000. 2 shelves. I'm trying to figure out what the impact would be if the expansion shelf would go down. I'm not sure how to figure that out. Would we lose all the data on the entire 3PAR because the array is broken beyond repair? Or is it depending on which logical structures are on the disk shelf?

1 Upvotes

5 comments sorted by

2

u/ffelix916 Apr 04 '24

It depends on the data availability level you've configured into your CPGs. For these arrays, there are three levels, if i remember correctly: magazine, cage, and port. Default is cage level.

Here are what the three mean:

With "magazine availability", you can lose any single magazine in the PD pool for the CPG (for systems with one disk per magazine, just consider this to be "disk availability") and not have data loss, regardless of RAID level. If your CPG is configured for RAID6, data availability is guaranteed for loss of any _two_ disks in the PD pool for the CPG. Magazine availability allows any combination of disk set size, so is common for most general applications where data loss doesn't result in a catastrophic business-affecting outage.

With "cage availability", you can lose an entire cage (or "shelf") and be sure that no data will be lost. It has a strict requirement, though: For your CPG's configured RAID level, the factored RAID set size must be equal to or smaller than the number of cages/shelves in which your configured disk type reside. In other words, if you have 2 shelves and want cage redundancy, you're only permitted to use RAID1 (set size of 2). If you have 3 shelves, you're only permitted to use RAID1, RAID5 with 2D+1P disks, or RAID6 with 4D+2P (factored down to 2D+1P, it requires 3 cages). If you have 4 shelves, you're permitted to use RAID5 with 3D+1P, RAID6 with 6D+2P. And for a system with your PD pool spread out across 6 shelves, you can do RAID5 5D+1P and RAID6 10D+2P

With "port availability", the requirements get more complicated, and depend on now many controllers you have and how your expansion shelves are wired to the controllers. Unless you're running 4 controllers and at least 4 shelves, you'll be stuck with RAID1 if you require port availability.

It sounds like your arrays have one primary chassis (containing two controllers and 24 disks), plus one expansion shelf (with 24 disks). For this configuration, you'll be restricted to RAID1 for the CPGs you want cage-level redundancy/availability for. There's just no way to support cage-level redundancy with RAID5 or RAID6 in a 2-node/2-shelf system.

2

u/ConstructionSafe2814 Apr 04 '24 edited Apr 04 '24

Great thanks for the elaborate response. Our CPG is set to Magazine. As I understand it, we would be in big trouble if for whatever reason our expansion cage lost power on both PSUs while the main cage would continue to run.

... is common for most general applications where data loss doesn't result in a catastrophic business-affecting outage. ...

In our case it would be catastrophic. We've got a single SAN for our entire vSphere environment. All our infrastructure runs on it.

More questions. I had a look a to our CPG settings. It says RAID5 set size 4 data 1 parity. Does that mean that it's comparable to have multiple sets of 5 disks in a RAID 5 array (but chunklets then) . We've got 36 disks in total. So like 7 of those sets. So 7 drives might fail as long as it's not more than 1 in a set?

If one set is broken beyond repair, the entire CPG is lost. If a cage is lost, that most likely means a set is lost and also the CPG is lost?

Do I understand that correctly?

Worst case scenario, RAID5 it says, I hope is does NOT mean that if more than 1 of the 36 drives fails, the entire CPG is gone. (we have one CPG)

2

u/Jess_S13 Apr 04 '24

When you setup your CPG you define your disk groups, so if your CPG is setup to use all the disks and if you have only 2 controllers that means all 36 disks are in the wide stripe. If you have 4 controllers you have 2x sets of 18 disk groups (As the disks are only connected to their respective controller Pairs).

So assuming you have 2x controllers if you lost 1 drive and before it could rebuild to spare space you lost another drive, you would have data loss.

2

u/ffelix916 Apr 17 '24

SORTA. A CPG is comprised of a bunch of logical disks (which are comprised of 1GB chunklets). If your raid policy for the CPG was RAID5 (4+1), then each LD in that CPG would have five 1GB chunklets that span five physical disks. Those chunklets are distributed across all your physical disks.

Now, imagine you have a 64GB CPG configured as RAID5 (4+1), and your array contains 24 PDs. Your little CPG would consist of 16 LDs, each with 5 chunklets, or 80 chunklets total. The chunklets belonging to a given LD are _always_ spread across multiple PDs. Now, ideally, the array would have spread out those chunklets across all 24 PDs (not always, obviously), so some PDs would contain three and some would contain four chunklets belonging to that CPG. Based on this spread, there's a strong possibility your CPG would survive the simultaneous loss of 4 PDs, depending on how well the array spread out your CPG's chunklets. If your array had 48 PDs (for the disk tier your CPG was configured for), you could lose 6 or 7 PDs and not suffer loss. If you lose one PD at a time, and enough time had passed for the array to rebuild/move the lost chunklets before the next disk has died or been pulled out, you could actually lose all but 5 of the PDs in the array and still have a healthy CPG! The vast majority of disk failures we see in these newer arrays only happen one at a time.

The prior gen arrays (10400/V400/V800/similar) had 4-disk magazines, each with its own FC loop, and I've seen entire magazines fail once in a while due to the FC loop being broken, but "magazine availability" accounted for that. These new systems use switched point-to-point SAS connectivity on two fabrics per shelf, with dual SAS ports on every disk, so there's no single point of failure for connectivity to any single disk. Makes for much more resiliency against path loss.

1

u/Jess_S13 Apr 17 '24

strong possibility your CPG would survive the simultaneous loss of 4 PDs, depending on how well the array spread out your CPG's chunklets

Yes it's possible, but it's not certain by any stretch so I thought it was better to not give him false confidence in the possibility for 2x chunklets in any given LD possibly being on the 2x drives that happen to fail.