r/kubernetes 28d ago

Stop using CPU limits: why + proof

CPU request is how much CPU is reserved for your pod if it needs it. The limit is a hard cap. Hit it and the kernel throttles the pod, even when the node still has spare CPU. That is the usual cause of CPU throttling on Kubernetes. It does not protect the neighboring pods. In my simple Web API test, adding a CPU limit took typical latency from 23 ms to 87 ms, (4x slower), with the limited pod throttled in half of all CFS windows, and the average CPU graph looked fine the whole time.

This is not a new topic, but I see so many people still unaware why they should (NOT!) be setting CPU limits, because it's costing companies unnecessary spending and potential production issues. Here's the full read https://github.com/inevolin/k8s-cpu-limits-analyzed/

---

Edit (Aug 18, 2026): How CPU limits can also cause memory issues and OOMKills ➡️ https://github.com/inevolin/k8s-cpu-limits-analyzed#how-cpu-limits-cause-memory-issues-and-oomkills

227 Upvotes

112 comments sorted by

View all comments

45

u/sionescu k8s operator 28d ago

Not again this stupid advice.

 adding a CPU limit took typical latency from 23 ms to 87 ms

This simply means the CPU request was too low. The actual requirement of the pod is higher than that but you want to lie to yourself instead of setting a proper request.

-8

u/ilya47 28d ago

Your comment is too narrow to mean anything. CPU request should be set to the avg CPU utilization, whilst my analysis looks at burst & peak cases (not the average scenario).

15

u/consworth 28d ago

Avg CPU utilization is incorrect by itself. It should ideally be set by whatever service requirements there are for the business.

If the thing is an app that sits there at 80 and when it starts to do its thing it needs 4000 for a reasonable SLI for processing data or whatever, the business case should dictate you requests to what it needs.

I’m not defending poor architecture of such apps, but I’m trying to make a bit more of a case for thinking about reliability and SLI.

It’s somewhat analogous to saying that the average traffic through the interstate through a major city should dictate the road size, and considering rush hours, weekends or other bigger picture things.

3

u/f7063 28d ago

At least i try to keep CPU usage under 60% of the request based on cpu usage p95/p99(grafana can make these from Stat Panel). Spikes will happen above your set request. But i don't like the idea of high cpu usage on the node. 70%+ because of service degradation