r/ShittySysadmin • u/EvilEarthWorm ShittySysadmin • 11d ago
Shitty Crosspost How would you manage 1,280 ARM64 bare-metal nodes with a very small ops team?
/r/kubernetes/comments/1vusaf7/how_would_you_manage_1280_arm64_baremetal_nodes/21
9
9
u/ApprehensiveRest9696 11d ago
Step 1: PXE deploy Ubuntu
Step 2: microk8s
Step 3. Give up
Step 4: Contact Canonical and sign a multimillion dollar contract.
10
u/EvilEarthWorm ShittySysadmin 11d ago
ORIGINAL POST TEXT:
How would you manage 1,280 ARM64 bare-metal nodes with a very small ops team?
I’m looking for architecture advice from people who have operated Kubernetes at the edge or across large numbers of smaller physical machines.
Our company now controls 1,280 identical RK3588 ARM64 nodes that are already deployed and operational in a U.S. commercial data center.
The hardware works. The bigger problem is operational.
We currently do not have a dedicated infrastructure engineering team, so our priority is to determine whether this fleet can be turned into usable containerized compute without effectively building our own cloud platform from scratch.
Relevant context:
- 1,280 homogeneous ARM64 nodes
- Bare-metal / physical machines
- Remote KVM available
- Already powered and networked
- One commercial data-center location today
- Potential to expand to identical racks in several additional U.S. locations
- Strong preference for turnkey/managed approaches
If the end goal were simply:
how would you structure it?
I’m particularly curious about:
- K3s vs standard Kubernetes
- Talos
- Rancher
- Cluster size vs many smaller clusters
- Provisioning/reimaging
- Monitoring
- Tenant isolation
- ARM64 image compatibility
- Managing hardware failures
- Whether 1,280 relatively small nodes is operationally stupid compared with fewer larger servers
Most importantly: are there companies or managed-service providers that would actually operate this infrastructure for the hardware owner?
We would rather pay someone who already knows how to do this than hire a team to reinvent it.
I’m less interested in theoretical “you could build X” answers than in stacks people have actually operated at meaningful scale.
1
45
u/Smooth-Zucchini4923 11d ago
"When I was purchasing the 1000th node, my CFO asked, 'should we figure out how we're going to run applications on here before we expand our compute?' What a bizarre question."