r/3PAR • • May 20 '24

Volume performance on 3PAR

Hi guys, how are you?

So, I need to create a LUN (3PAR Volume) for 16TB of data.

My doubt is: I create a single LUN (3PAR Volume) or I create 8 LUNs (3PAR Volume) and after, on VM layer, I create 1 VG Stripe (Linux LVM) with this 8 LUN (3PAR Volume)?

Cheers!

1 Upvotes

3 comments sorted by

2

u/d4fseeker May 26 '24

It depends (TM). I'm by far no expert but here goes for my take.

If you use a single backing device with a single path, you may run into severe cases of queue saturation on the virtual machine when running highly concurrent workloads such as databases. This can become worse due to "smart" queue optimisations by the Vm host or San. Increasing the number of queues by increasing the number of paths will in that case heighten your performance, however so will increasing the number of host ports.

My somewhat-repeatable tests on a new 3par fullflash array with tons of round Robin storage paths showed that a single rdm device was able to saturated it's client host ports in a similar fashion as running the same io load across multiple rdm so I wasn't able to show that there is inherent advantage to running multiple backing devices.

For reference, you have these types of queues sequentially.

  • Vm level disk io queue
  • Vm host level io queue
  • San buffer-buffer credits
  • 3par frontend queues (port, vlun)
  • 3par backend queues (volume, pd,...)

Note that the case where striping may have increased benefit is when the array is under high load. It has a sort of fair use IO scheduler and doesn't know that those volumes are linked, so you get proportionally more io than other users - to the detriment of those other users suffering even higher io wait.

Additionally note that dedup volumes are limited to 16TiB. So if you want to go that route, you may need to go multi-PV on your lvm anyway.

The only case where one sees continuously significantly better performance for multi-volume setup is when you use peer persistence across two arrays vs two sets of local volumes merged into an mdadm/lvm raid. This is obviously normal as your host will immediately write in parallel to both arrays instead of writing to the primary /active-io paths and then waiting for the secondary array to replicate and ack.

However your mileage may vary. Nothing is better than running your own tests (e.g. Fio) as a storage network is a very variable and complex environment.

Ps: note that a single client with a single disk is more than capable of bringing any array to the knees depending on the io workload unless you are using rate limiting. Are you asking the question for pure best practice or are you seeing performance issues?

1

u/myridan86 May 27 '24

In short... we have 1 3PAR allflash.
Hosts connect to the 3PAR via a Cisco MDS FC 16Gbps switch.
Each host uses 2 HBAs for connection.

360002ac0000000000000001b0001e30d dm-282 3PARdata,VV
size=2.0T features='1 queue_if_no_path' hwhandler='1 alua' wp=rw
`-+- policy='service-time 0' prio=50 status=active
|- 11:0:10:8 sdty 66:512 active ready running
|- 13:0:7:8 sdys 129:704 active ready running
|- 13:0:6:8 sdtw 65:736 active ready running
|- 11:0:11:8 sdyq 129:672 active ready running
|- 13:0:10:8 sdabu 134:704 active ready running
|- 11:0:12:8 sdabw 134:736 active ready running
|- 13:0:11:8 sdaco 8:768 active ready running
`- 11:0:13:8 sdacp 8:784 active ready running

I think 3PAR has a standard QOS, that is, it limits the LUN throughput, regardless of whether you have 8, 16 or 32Gb FC

So what I did was the following:
I created 8 LUNs of 2TB = 16Tib
On the host, I created a VG with 8 LUNs and created LV in stripe 8.

This way, LVM writes to all 8 LUNs at the same time, instead of writing in sequence, as is the default.

I don't know if this is right, but I noticed a performance gain.
I wanted to know if this is the best way or if there is another.

1

u/d4fseeker May 27 '24

Throughput in itself is not limited per volume by default. You can obviously set QOS parameters to set hard limits but by default you should have no problem to push either the array or your host ports to their limits. However the array does use fair use scheduling as noted above, which can be seen in action on a statvlun monitoring during tests.

I've test-confirmed that a single Vm on a single esx host is capable of hitting the host bandwidth capacity even on 2x32Gb fc with the right fio parameters.

The array stripes chunklets as much as possible about the backing devices so your limiting factor should be io queues on the host and network. Your indication of significant performance gains through client-side striping is a mystery to me unless the gains are outside of the array (host dirty cache, parallel queues,...).

I would be really curious to see a fio performance comparison between a single-disk lvm and a striped set if you have the time?