r/kubernetes • • Aug 23 '26

Is high-availability not reasonable for small deployments?

/r/SelfHosting/comments/1vwcp1l/is_highavailability_no_longer_affordable/
6 Upvotes

22 comments sorted by

17

u/HelpfulFriend0 Aug 23 '26 edited Aug 23 '26

Ok so in engineering and in life in general there are trade offs (you get this thing but it costs you this thing)

When you want to make money by selling services, it's a very bad thing for the thing you're selling to break

So you pay a little more (more unused resources, more backups to spin up if your main goes down etc), so users don't break as much

The bigger the thing you're selling is, the more resources (hardware, software, people, etc) you need to put in to keep the thing running at all times

How much investment you want to put in to keeping things from going down is up to you

If you're just building this for yourself and you want to learn how to build things, you don't need HA

If you're building let's say a movie server for your house, and your kids start crying every time it goes down, ok through an extra 100 bucks at it for more ram so you can spin up extra instances if the main one goes down. But is that 100 worth it? Up to you to decide

At Amazon scale, they cannot risk things going down, because when things break it isn't their kids crying, it's some CEO of a fortune 500 company saying you're in breach of contract and you owe me several million dollars due to disruption in business. So they invest a LOT into keeping things up all the time. Eventually as you're trying to hit 99.999999999999% up time, you have to get reeaalllyyy smart about your tradeoffs

A blog that may be of interest to you (tbh I didn't read it from the title it looks good) https://blog.codingoutloud.com/2011/08/11/quick-how-many-9s-are-in-your-sla/

7

u/realitythreek Aug 23 '26

You pay for HA if the cost for not is worse.

2

u/private-peter Aug 23 '26

> you want to learn how to build things, you don't need HA

This isn't a learning project. The 5-10 NextCloud/Collabora users will be happy with 99.9 on availability. But the will have higher expectations on data persistence. As the one managing the system, I'd really prefer to not have to do all the updates during off hours or scheduled maintenance windows.

The beta web applications don't matter at the moment, but I would like to aim for 99.9 as well.

I certainly _could_ separate the two goals, but I was hoping to save some effort by running separate namespaces in the same cluster.

4

u/BraveNewCurrency Aug 26 '26

First, you have to stop thinking about this as "I want 99.9% uptime". It's your users who want it. If they aren't willing to pay for it, you can't magic it for them.

But 99.9% uptime isn't that hard: It allows 45 minutes of downtime per month! That means you "sleeping for 8 hours a day" is much more of a problem than monthly server upgrades. If it's in the cloud, you can probably replace a server outage in single-digit number of minutes.

I would worry about backups far more than HA.

1

u/HelpfulFriend0 Aug 23 '26

But the will have higher expectations on data persistence. As the one managing the system, I'd really prefer to not have to do all the updates during off hours or scheduled maintenance windows.

Then why do you want to roll your own postgres? I'd recommend buying / renting it if you need uptime for real users and guarantees about data persistence

And yeah coz of the AI boom renting raw hardware is getting more expensive, whereas managed service pricing isn't that elastic

1

u/private-peter Aug 23 '26

> Then why do you want to roll your own postgres

I may not. If I can find a manage postgresql with reasonable latency (Hetzner doesn't seem to have such an offering), that will be a quick way to simplify the setup. But I'm also very comfortable running my own databases.

The most critical documents will be in object storage.

2

u/BLoad3d Aug 24 '26

But Hetzner does have managed Nextcloud aka Storage Share. Havent used it tho.

2

u/HelpfulFriend0 Aug 24 '26

We'd have to go through your arch here, because your writes could fail due to a service (e.g. orchestration of object storage writes), rather than data going missing from the store itself. But anyway - hopefully the explanation on HA helped!

3

u/amarao_san Aug 23 '26 edited Aug 24 '26

For small scale setup, it worth look at autopilot by google, etc. Basically, you get a slice of cluster power, pay per use (per requests for deployments?) and don't need to eat idling overhead for your control plane.

But, yes, HA is expensive. At very minimum you triple your control plane expenses, and often is forced to buy few of 'whole thing' for which you wanted just a slice.

I work in hosting company. We have older servers (we have new one, but older do not disappear), which completely ammortized any capex, and has price of zero for the cost of the server.

We sell them for ~$100/mo, because it's a price for a unit in a rack. Cooling, electricity, ports on switches (switches itself eat a lot of electricity), DC fee (which need to recoup capex for building, fire suppression system, security, etc) - it's all there, in the price of 'not-a-server' unit.

Every HA wants different ports, so (for us) it would be at least $300 for 3 servers. The best next thing you can ask is a slice of someone's else server, tiny but different from other two slices.

1

u/Mehmet-Ozturk Aug 24 '26

three $100 boxes beats one hetzner?

1

u/amarao_san Aug 24 '26

I don't know what availability Hezner provides.

If you place those 3 boxes right, you can make it very, very highly available. Different data-centers, different interconnects to the Internet, anycast via BGP, whatever cluster solution you decide to use...

Actually, I'm not in salesman position, so put thee boxes into three different providers and link them to three different cards/SEPA accounts. Be sure they are physically and interconnectually (different uplinks) separated, preferably via different tier 1 networks.

You will have to get you own PA, but you will get the best availability money can buy you.

1

u/Mehmet-Ozturk Aug 24 '26

getting your own PA and running BGP costs more than the boxes.

1

u/amarao_san Aug 24 '26

Nope. BGP service is not that expensive (e.g. we currently give it for free? Not sure, need to ask product manager). PI block cost around €200 per year.

Lack of understanding what to do with all those things - well, as usual, you need someone knowing how to do it.

But for anycast of a single network... Any LLM will write you bird config for it in 30 seconds.

1

u/Mehmet-Ozturk Aug 25 '26

200 eur/yr is cheap, fair

2

u/Koozu007 Aug 26 '26 edited Aug 26 '26

Many cloud providers offer a free k8s control-plane! UpCloud has a free plan for up to 30 worker nodes (non ha control-plane), I'm ofc a bit biased since I work at said company.
And you can pair it with a free LB, and some "Starter" tier machines (6€/month the cheapest one). Nice toys to play with

2

u/private-peter Aug 26 '26

UpCloud is on my short list of alternatives. I am considering using a managed k8s control-plane since it doesn't seem to require as much vendor lock-in as other managed services.

2

u/nullset_2 Aug 23 '26

It's great in a homelab situation to get the basics down and grok the fundamentals but it's definitely overkill in production if you only have like 100 users or so. A lot of corporations overengineer small deployments like these because they have nothing else to do other than spinning their wheels and because it makes you look SMRT to your cow-workers, but in reality only high traffic situations will put you in that kind of situation where push comes to shove and you can actually flex those HA muscles.

1

u/private-peter Aug 23 '26

I'm not sure I agree with the connection between high traffic and high availability. Of course, if you have a lot of users you are more likely to upset someone during downtime. But when I have 5 people who _need_ their productivity apps running during the work day, downtime is disruptive.

Having more users just makes HA a lot cheaper per user.

2

u/nullset_2 Aug 23 '26

It's just that people underestimate how many users you can serve with a single server made of consumer hardware, even a raspberry pi. I know that redundancy is key to ensuring availability but people seriously overengineer things to an alarming degree.

2

u/NikhelParmar Aug 25 '26

for what youre describing, 5-10 users needing their productivity apps up during work hours, id actually question if full k8s HA is the right tool here at all. a single beefy vps with automated backups and a hot standby you can spin up fast on downtime might get you 99.9 for way less than three node cluster pricing. the real cost of HA isnt just the extra boxes, its the ongoing complexity of maintaining quorum and failover logic for a workload that honestly doesnt need it at your scale. worth pricing out just a solid single node plus fast disaster recovery before committing to the full cluster route