For the love of god - care about other pods on the node, especially in a multi-tenant setup.
Sorry for the cheeky response.
CPU Limits have a place, you don't want a bad change for 1 deployment object affect all neighbors by taking all the CPU. You need to be able to constrain the blast radius. This doc gives me strong AI vibes. Setting CPU limits isn't free. You still need to care about how the programming language that you use discovers those limits, and correctly handles them. For e.g. if you spin up a 100 Java threads, but only have 1 cpu as the limit, that's bad design.
While I agree that CPU limits tend to make your performance worse, I don't think the delivery of the post is all too convincing (and is pretty heavy on the LLM-isms that it's putting me off from reading).
It mentions that a cpu request is a guarantee, but how is that enforced? If I have 32 pods running on a 32 core machine, each with 1cpu requested, what stops one of those pods using an unfair share? I assume we just rely on the Linux scheduler. If I have 16 pods with 1cpu and 1 pod with 16cpu, does the Linux scheduler make sure to give the 16cpu pod more time? Or are we back to using cgroups.
Essentially, they avoided CFS quota throttling by assigning exclusive CPU cores via cpusets. That sacrifices some burstability and packing efficiency in exchange for stronger, more predictable CPU isolation.
This always sounds good on paper and this is very common lore, but then when you get into production escalations a very common problem is a lot of software depends on limits for autoconfiguration of thread pools and even several runtimes (go and java for example, at least .net is mentioned in the article), yes you can usually set them with a flag but people have to know this, communicate it, enforce it. Basically replace adhoc what limits is doing for you automatically configuration wise
So this just all assumes you have a setup where all teams communicate the necessary information perfectly.. what happens in practice is workloads degrade at edge cases because there are 256 threads running for a thread pool instead of 4.
I would expect the impact of cpu limits to be different between K8s providers. I have only used memory limits. If I had a pod that tended to be very cpu intensive, I would schedule it on it's own node group.
Kubernetes CPU limits can also cause memory issues and OOMKills. It looks like a memory leak, the usual fix is a bigger memory limit, but the actual cause is the CPU limit. Here's some simple proof: https://github.com/inevolin/k8s-cpu-limits-analyzed#how-cpu-...
15 comments
[ 0.24 ms ] story [ 43.2 ms ] threadSorry for the cheeky response.
CPU Limits have a place, you don't want a bad change for 1 deployment object affect all neighbors by taking all the CPU. You need to be able to constrain the blast radius. This doc gives me strong AI vibes. Setting CPU limits isn't free. You still need to care about how the programming language that you use discovers those limits, and correctly handles them. For e.g. if you spin up a 100 Java threads, but only have 1 cpu as the limit, that's bad design.
It was already the case in 2018.
Also, no mention of the scheduler overhead. And the maintenance overhead is the worst.
It mentions that a cpu request is a guarantee, but how is that enforced? If I have 32 pods running on a 32 core machine, each with 1cpu requested, what stops one of those pods using an unfair share? I assume we just rely on the Linux scheduler. If I have 16 pods with 1cpu and 1 pod with 16cpu, does the Linux scheduler make sure to give the 16cpu pod more time? Or are we back to using cgroups.
Essentially, they avoided CFS quota throttling by assigning exclusive CPU cores via cpusets. That sacrifices some burstability and packing efficiency in exchange for stronger, more predictable CPU isolation.
So this just all assumes you have a setup where all teams communicate the necessary information perfectly.. what happens in practice is workloads degrade at edge cases because there are 256 threads running for a thread pool instead of 4.
(The title of this was also stolen for this HN post, although the GitHub repo makes no mention of it...)
It's the classic "developer can't optimize their stuff, so they ask for infinite resources to cover their mistakes" farce.
No, keep the CPU limits. They exist for a reason. Fix your app.