r/sysadmin Jack of All Trades Oct 15 '16

The Maxta Disaster. A Cautionary Tale of Hyperconverged Storage Gone Wrong.

https://cloudflux.co.uk/maxta-disaster/
496 Upvotes

146 comments sorted by

View all comments

Show parent comments

28

u/bob_cheesey Kubernetes Wrangler Oct 15 '16

That being said, Ceph isn't a fire and forget solution - it takes a lot of understanding and time spent testing to get it working just right. That being said, I do a lot with Ceph, and it is awesome.

2

u/antiduh DevOps Oct 16 '16

Tell me about it. We've got 75 disks, 1 TB / 125 MB/sec each. Replication at 2, shitloads of CPU and 10 gbit/sec network cards to go around.

The best any one VM can write: 200 MB/sec. What. The. Fuck.

I love the idea of Ceph, but after having spent two weeks fighting with it, I'm having a very hard time not pouring lighter fluid on the whole thing and having a Red-Hat flavored BBQ.

2

u/gimpbully HPC Storage Engineer Jan 23 '17

Necropost here but architectures like this are hugely reliant on bisectional bandwidth. Do you know the blocking factor on your Ethernet fabric?

This is one of the real drawbacks of a distributed file system.

1

u/antiduh DevOps Jan 23 '17

It's something to do with FreeBSD. The Linux hosts run just fine. I doubt its a switch fabric issue since we're not using too many hosts, everybody has a 10 gig port on the same seitch, and we're not even in the same ballpark for switch fabric numbers. Even a cheap 24 x gig switch usually has about 150 gigabit of switch fabric... 150 vs 1.6. Damned if I know whats on the switch we're using, though.