Are Kubernetes resource limits still misunderstood?
CPU limits cause throttling even when the node is idle. Memory limits cause OOM kills. Many teams set both identically to requests without thinking about which they need.
CPU limits cause throttling even when the node is idle. Memory limits cause OOM kills. Many teams set both identically to requests without thinking about which they need.
Memory limits: yes, always. CPU limits: rarely. Requests do most of the scheduling work you actually want.
Unless you run multi-tenant clusters. Then CPU limits protect neighbours, throttling or not.
Look at the container_cpu_cfs_throttled metric before arguing. Most people have never checked it.
Some numbers from a cluster I ran last quarter, since this debate usually happens without data.
We had ~300 services with CPU limits = requests. I removed CPU limits from 40 latency-sensitive services as an experiment (kept memory limits).
- p99 latency dropped 18–35% on most of them - Node CPU utilization went up ~6% - Zero noisy-neighbor incidents over 8 weeks - Two services did go wild during a traffic spike, but requests kept them from starving others because CFS shares still apply
The throttling was the big one. Several services showed 20–40% of periods throttled while using, on average, less than half their limit. Short bursts on many threads hit the quota inside a 100ms window even though the average looked fine.
sum(rate(container_cpu_cfs_throttled_periods_total[5m])) by (pod) / sum(rate(container_cpu_cfs_periods_total[5m])) by (pod)
If that ratio is above a few percent on a latency-sensitive service, the limit is probably hurting you more than it's protecting anyone.