Skip to content
10 min read

Kubernetes Memory QoS: What Changes After the Upgrade

Kubernetes 1.37 changes Memory QoS defaults. Learn how to preserve intended behavior, evaluate throttling and reservation separately, and judge a canary by service health.

Antonio J. del Águila

Knaisoma

Kubernetes 1.37 enables the Memory QoS feature gate by default, but the default kubelet configuration enables neither early throttling nor memory reservation. Each behavior has its own setting. Teams adopting the feature for the first time must choose which behavior they need; teams that enabled it during alpha must check whether they relied on a default that has changed.

The practical question is whether the resulting workloads serve users better. Throttling can give the kernel an opportunity to reclaim memory before a hard limit is reached, but it can also introduce sustained latency before an eventual OOM kill. Reservation protects workloads from reclaim while reducing the memory available to their neighbors. Fewer restarts alone do not demonstrate a healthier service, so both settings need a workload-specific acceptance test.

What the 1.37 upgrade actually writes

The release post for Kubernetes v1.37 is explicit. Turning the feature on by default is described as safe because the default kubelet configuration does not enable memory throttling or memory reservation, so no memory.high, memory.min or memory.low values are written to cgroups unless you configure them. Two fields do the configuring. Setting memoryThrottlingFactor to a value between 0 and 1 enables memory.high throttling for Burstable and BestEffort containers. Setting memoryReservationPolicy to TieredReservation enables protection through memory.min and memory.low. Their defaults are null and None.

The interesting part is the history. In earlier alpha releases, memoryThrottlingFactor defaulted to 0.9, so enabling the gate made the kubelet set memory.high on containers. The v1.36 post on tiered protection says exactly that: enabling the gate turns on throttling with a default factor of 0.9. Version 1.37 changed the default to null because an automatic memory.high could throttle workloads that had been running unthrottled, and the project wanted an upgrade to leave runtime behavior alone.

Two consequences follow for anyone who touched the feature before. If your kubelet configuration file already contains an explicit memoryThrottlingFactor, that value is preserved. If it does not, and you had been relying on the old default, the 1.37 kubelet stops setting memory.high. Throttling you thought you had quietly disappears on upgrade, and nothing fails to tell you.

The second consequence is about reading advice. A vendor write-up published on September 2 still describes 0.9 as the default, whereas the September 14 upstream post specifies null. Verify the configuration and effective cgroup settings on your nodes. A container cgroup whose memory.high reads max has no throttle configured at that level; ancestor cgroups and node-wide memory pressure still matter.

Lever one: throttling can introduce latency before an OOM

For a Burstable container, the kubelet computes the throttle point from the request and limit, and the Kubernetes documentation gives the formula: memory.high = requests + memoryThrottlingFactor * (limits - requests). Its worked example is a 256 MiB request with a 1 GiB limit at a factor of 0.9, which yields roughly 947 MiB. Guaranteed containers get no memory.high, because their requests equal their limits. A Burstable container with no limit has node allocatable memory substituted for the limit, which puts its threshold most of the way up the node. A BestEffort container is treated as having a zero request and the same substitution.

What the kernel does at that threshold matters more than the arithmetic. The cgroup v2 documentation says that when usage exceeds memory.high, the cgroup’s processes are throttled and put under heavy reclaim pressure, and that crossing this threshold does not itself invoke the OOM killer. The high boundary can be exceeded. The hard limit, memory.max, still does: when usage reaches it and cannot be reduced, the OOM killer runs. Kubernetes leaves memory.max where it was.

Here is an illustrative scenario, constructed for this article and not drawn from any client. A Burstable API container requests 256 MiB, is limited to 1 GiB, and has a slow leak. Without throttling, the leak reaches the limit after some hours, the container is killed and restarted, and the restart counter and an OOMKilled reason make the event easy to find. With a factor of 0.9, the container reaches about 947 MiB and its allocations begin to stall while the kernel reclaims aggressively. If the leak outpaces reclaim, it still reaches 1 GiB and is killed, only later and after a period of degraded latency. While reclaim keeps allocations moving, the container may remain degraded for an extended period. Throttling does not repair the leak or guarantee that the container will avoid an OOM.

A liveness probe may keep passing while user requests slow down, depending on what its endpoint checks and its configured thresholds. Kubernetes uses failed liveness probes to trigger container restarts; that recovery mechanism does not replace request-level monitoring. The scenario illustrates a risk implied by the kernel semantics, not a measured result: dashboards centered on restarts can miss sustained degradation. Throttling can be the right trade for a batch worker that tolerates slowness, or for a node where one runaway container was previously taking its neighbors down. It is a poor trade for a latency-sensitive service whose requests are set carelessly, because the formula builds the threshold from the request, so an unrealistic request produces an unrealistic threshold.

Lever two: reservation spends node headroom

TieredReservation protects memory instead of limiting it. According to the v1.36 post, Guaranteed pods receive hard protection through memory.min, which protects usage within the effective minimum boundary from reclaim. That boundary is constrained by ancestor cgroups, and overcommitting protection can cause OOMs when no unprotected reclaimable memory remains. Burstable pods receive soft protection through memory.low, which the kernel avoids reclaiming under normal pressure but may reclaim to prevent a system-wide OOM. BestEffort pods receive neither.

The tiering exists because the earlier behavior was blunt. The same post gives a node with 8 GiB of memory and 7 GiB of Burstable requests: before tiering, all of that would have been locked as memory.min, leaving little room for the kernel, system daemons or BestEffort workloads. With tiering, only Guaranteed pods are hard-protected, so the amount of unreclaimable memory is smaller and the risk of node-level OOM kills is lower.

Tiering does not remove the cost, and the 1.37 post is candid about the remaining limits. The policy is node-wide: every Guaranteed pod gets memory.min and every Burstable pod gets memory.low, with no per-pod opt-in or opt-out, so a node that mixes workloads needing hard reservation with workloads that should stay reclaimable must pick one policy for all of them. Hard reservation also covers everything charged to the container’s cgroup, including page cache, so a pod that reads large files can hold memory the kernel would otherwise reclaim to serve its neighbors. SIG Node is tracking both in kubernetes/kubernetes#140246, which is still open.

The v1.36 post suggests a sensible order: enable throttling first, observe workload behavior, and opt into reservation when the node has enough headroom. It also introduces two alpha-stability kubelet metrics, kubelet_memory_qos_node_memory_min_bytes and kubelet_memory_qos_node_memory_low_bytes, whose totals are useful for capacity planning. If the memory.min total creeps toward the node’s physical memory, hard reservation is getting tight.

How much benefit to expect

The release posts describe mechanism and rollout, not measured benefit. The only published numbers I found come from a different implementation: Alibaba Cloud’s documentation for container memory QoS reports a Redis benchmark under memory overcommitment where average latency fell from 51.32 ms to 47.25 ms and throughput rose from 149.0 to 161.9 MB/s. That setup uses its own controller, ack-koordinator, with parameters that include a watermark mechanism and some dependence on Alinux kernel features, and the page itself says results depend on cluster configuration and workload. It is evidence that the idea can help under overcommitment, not a forecast for upstream Kubernetes on your nodes.

The cited benchmarks do not establish the benefit for upstream Kubernetes on your workloads. Use them as context for a hypothesis, then test that hypothesis with a representative canary before expanding the change.

Make the quiet failure visible first

Throttling can introduce sustained degradation that restart counts alone will miss, so confirm that instrumentation covers both before changing the setting. The cgroup interface already exposes what you need. In memory.events, the high counter records how many times the cgroup’s processes were throttled and sent into direct reclaim, max counts how often usage was about to exceed the hard limit, and oom_kill counts processes killed by any OOM killer, all as documented in the kernel guide. In memory.pressure, the pressure stall information interface reports the share of time in which some tasks, or all non-idle tasks, were stalled on memory, with averages over 10, 60 and 300 seconds plus a running total.

A useful canary records restart and oom_kill counts, request latency percentiles, error rates and throughput before any change. Capture memory pressure for the affected workload cgroups as well as the nodes; node-level averages alone can hide a struggling container. For batch workloads, record queue growth and completion times. Compare representative traffic or jobs against an unchanged pool where practical, and change one setting at a time so the result remains interpretable.

After enabling a throttling factor on one pool, correlate increases in the memory.events high counter and the some and full pressure values with those service-level signals. A rising high counter establishes that throttling occurred; it does not by itself establish that users benefited or suffered. Keep OOM and restart monitoring because both remain possible.

Define the acceptance rule before starting: retain the change only if latency and error rates stay within the workload’s service objectives, batch jobs meet their completion targets, and neighboring workloads remain healthy under representative peak load. A reduction in OOM kills that comes with missed service objectives fails that test. Name the rollback owner and observation window, and compare rates over comparable periods rather than treating a short interval without a rare OOM as proof of improvement.

Before either experiment, verify cgroup v2, the documented kernel and runtime prerequisites, and each pool’s existing configuration. Preserve an explicit throttling factor if you intend to retain earlier behavior. Then evaluate throttling and reservation independently: reservation can be enabled without throttling, as the v1.37 release post demonstrates. A throttle-first rollout is a way to isolate changes, not a dependency between the settings.

First verify node prerequisites and preserve any explicit factor needed for existing throttling. Decide independently whether to test throttling and whether to test tiered reservation. Test one change at a time, and retain it only if service objectives and neighboring workload health remain acceptable.

flowchart TD
  A["Verify prerequisites;<br/>preserve intended settings"]
  A --> B["Throttling choice:<br/>off or canary a factor"]
  B --> C["Independent protection choice:<br/>None or canary TieredReservation"]
  C --> D["Test one change at a time;<br/>accept against service objectives"]
Preserve intended behavior, then evaluate each Memory QoS setting independently

Rollback is not a flag flip

One operational detail is easy to miss because it only matters on the day you want out. The release post says that to disable the feature after upgrading, you set the gate to false and ensure a compatible kubelet configuration, and the kubelet rejects the configuration if memoryThrottlingFactor is set to anything other than the former default of 0.9, or if memoryReservationPolicy is TieredReservation. A team that enabled both fields and wants to back out has to remove those fields in the same change as the gate, across every node pool. When the gate is off, the kubelet resets stale protection at startup, and stale memory.high values on containers are reset to max on paths such as restart or resize, which means the old limits clear gradually rather than all at once.

A rollout plan should therefore include the rollback configuration, reviewed in advance, with the node pools it applies to named.

What this means for a platform team

The gate flipping to on is not the decision. The decision is who owns the two fields, and the honest answer is the same team that owns requests and limits, because both levers are derived from them. A throttle point built from a placeholder request and a reservation built from an inflated one will each behave according to numbers nobody chose carefully.

Before the next upgrade window, it is worth confirming what each node pool actually has configured and checking one container cgroup per pool, so that the platform’s real state replaces the assumption that the new default did something. Then choose each setting per pool against the service objectives and capacity constraints of the workloads running there.

Memory settings are easy to change and hard to evaluate, which is why they tend to be left alone until an incident forces the question. We help platform teams audit kubelet configuration and resource requests against real workloads, design canaries with the right memory signals, and plan upgrade and rollback paths for node-level features like this one. Talk with us about your Kubernetes memory and upgrade planning.

Kubernetes Operations Reliability Architecture
Share:

Stay updated

Get insights on engineering transformation delivered to your inbox.

Newsletter coming soon.