The Kubernetes resource limits decision record: why the QoS class you assigned determines your OOMKill blast radius and your CPU throttling latency floor

The resource limits model — pods running as BestEffort with no requests or limits because no one had time to measure actual usage when the initial Kubernetes manifests were written, CPU limits set equal to CPU requests to enforce "predictable" resource usage without understanding that the Linux CFS quota mechanism converts efficient bursty CPU consumption into a wall-clock throttle that adds hundreds of milliseconds to tail-latency requests, and memory requests set at average usage rather than peak usage to maximize pod density per node — are Kubernetes resource configuration decisions that are almost never made explicitly. They emerge from a Kubernetes adoption where the first concern was getting services running and the resource configuration was deferred for a future optimization sprint that arrives only after the first major incident; from a convention that "predictable limits" are safer than burst-capable configurations, established by an operator who understood the Guaranteed QoS class but not its CPU throttling implications for bursty workloads; and from a capacity planning model that conflates scheduled capacity (requests) with actual usage, producing a node over-subscription surface that is invisible to the scheduler and only visible when a correlated peak usage event strikes multiple pods simultaneously. Three failure patterns: the developer productivity SaaS that ran all 12 services with no resource limits and watched a memory-leaking report aggregation pod consume 4.2 GB on an 8 GB node during a batch run, forcing the OOM killer to terminate system processes and cascading two of three nodes into NotReady state in a failure that required 80 minutes to recover; the fintech SaaS that set CPU limits equal to CPU requests for predictability and spent three weeks debugging what appeared to be a race condition in its transaction validation service before discovering that CFS throttling was consuming 65% of wall-clock time on its most complex requests and producing a 20x gap between P50 and P99 latency; and the developer tools SaaS that set memory requests at 50% of measured average usage to maximize pod density, scheduled eight indexing pods onto two nodes, and had all eight OOMKilled simultaneously when a large customer's repository reindex drove all eight pods to their peak memory usage at the same time.

A 36-person developer productivity SaaS company built a sprint analytics and team performance metrics platform for engineering managers — cycle time tracking, deployment frequency measurement, PR review latency, and bottleneck identification for software teams. Their backend ran 12 services on Kubernetes, deployed on a three-node EKS cluster with 8 GB of memory per node. The Kubernetes manifests had been written by the infrastructure engineer who initially set up the cluster, who had annotated the resource fields with # TODO: tune these and left the requests and limits unset. The convention had propagated: each service added over the following year followed the same template with the same unset fields. All 12 services ran as BestEffort pods — no CPU or memory requests, no CPU or memory limits.

Under normal operating conditions, the BestEffort configuration was not the source of any incident. The cluster had sufficient memory across its three nodes to accommodate the services' typical memory usage, and the Linux OOM killer had never needed to intervene. The team was aware that the resources were unset but had no specific operational signal that this was a problem — services started, ran, and responded to requests, and the cluster's overall memory utilization dashboard showed 45% average usage, well below any danger threshold.

In the third quarter of the product's second year, a product manager wrote a script to regenerate quarterly reports for all active customers before a board presentation. The script iterated over 847 customer accounts and triggered the report aggregation service — a Python process that pulled sprint metrics from the primary PostgreSQL database and assembled them into a data structure for chart generation — for each account in sequence. The aggregation service had a memory leak: intermediate result sets for in-progress calculations were retained in a dictionary that was never garbage-collected between requests. Under normal single-request usage, the leak was invisible; the process served one request, accumulated a few hundred kilobytes of unreleased data, and waited for the next request. Under the bulk report generation, the process was serving requests continuously with no idle period between them: the dictionary grew with each request, never cleared, accumulating the intermediate data for 847 sequential report generations.

After 14 minutes and approximately 400 accounts, the report aggregation pod's memory usage had grown from its typical 180 MB to 4.2 GB. The node hosting the aggregation pod had 8 GB of total memory, with the remaining 11 services consuming approximately 3.6 GB collectively. Total memory usage on the node was 7.8 GB — 97% of the node's physical capacity. The Linux kernel's OOM killer began evaluating processes to terminate. The OOM killer calculates an oom_score for each process based on the ratio of its current memory usage to its total addressable memory space, with adjustments for process criticality; BestEffort containers have the highest oom_score and are killed first. The aggregation pod's process had the highest oom_score given its 4.2 GB usage. The OOM killer terminated the aggregation pod's process, freeing 4.2 GB. The node returned to safe memory utilization. Kubernetes detected the pod termination, classified it as an OOMKill, and restarted the pod.

But the aggregation service had been mid-request during the OOMKill, and the in-flight bulk report generation job had been aborted. The product manager restarted the script from where it had left off. Within 11 minutes, the aggregation pod reached 4.2 GB again. This time, the OOM killer's intervention was slower: other services on the node had increased their memory usage in the interval between the first and second OOMKill, and by the time the OOM killer was triggered again, three other service pods had elevated memory usage from processing requests that had accumulated while the aggregation pod dominated the node's memory. The OOM killer killed the aggregation pod, then killed two other service pods to free sufficient memory, then — critically — the OOM killer's memory reclamation was insufficient to restore the kernel's memory allocation zones to a usable state before new allocations were attempted. The kernel began terminating kube-proxy and kubelet system processes to reclaim memory. Once kubelet was terminated, the node's status reporting to the Kubernetes control plane stopped, and the control plane marked the node as NotReady.

Kubernetes attempted to reschedule the evicted pods from the NotReady node to the remaining two nodes. The two healthy nodes each had 8 GB of memory, with existing pods consuming approximately 5 GB each. The rescheduled pods — without any resource requests, as BestEffort — were placed based on available memory headroom, which Kubernetes estimated at 3 GB per node. In practice, the memory usage of the existing pods on those nodes was higher than 5 GB; the Kubernetes scheduler had no visibility into actual usage, only into the scheduled (zero) requests. Within 4 minutes of rescheduling, the two remaining nodes also entered memory pressure, and within 8 minutes a second node entered NotReady state. The cluster was effectively reduced to one functional node hosting five services. Recovery required spinning up replacement nodes (14 minutes for EKS node provisioning), waiting for pods to reschedule and start, and re-establishing service-to-service connectivity — 80 minutes total. Connect this failure to the capacity planning decision record: the cluster capacity planning model had measured total memory utilization against total cluster capacity and concluded the cluster was at 45% utilization with comfortable headroom; the model was correct at the aggregate level and completely incorrect at the individual node level, because with BestEffort pods, a single runaway process could consume 100% of a single node's memory without the aggregate cluster utilization metric showing any warning; resource requests are the mechanism by which capacity planning connects to scheduling reality — without requests, the scheduler distributes pods based on heuristics rather than committed capacity, and the capacity model becomes a fiction that describes what is scheduled, not what is used.

A 44-person fintech SaaS company built a payment compliance and reconciliation platform for neo-banks — transaction validation against rule sets, regulatory reporting automation, and audit-trail generation for payment processors. Their backend ran 18 Go services on Kubernetes. The infrastructure lead had applied a uniform resource configuration policy to all services: CPU request equal to CPU limit, memory request equal to memory limit, ensuring that every pod ran as Guaranteed QoS. The policy was documented in the internal Kubernetes standards as "predictable resource allocation prevents neighbor effects and makes capacity planning straightforward." The CPU limit for each service had been set at 0.5 CPU cores — estimated from the service's average CPU usage at the time the configuration was written, rounded up for headroom.

The transaction validation service — the component that evaluated each transaction against a set of compliance rules before authorizing it — had two distinct request types: simple transactions (standard card payments against a small rule set) that completed their validation in 15–40 milliseconds of CPU time, and complex transactions (cross-border payments, high-value transfers, multi-leg settlements) that evaluated against a larger rule set and completed in 350–600 milliseconds of CPU time. Complex transactions represented approximately 4% of request volume but carried higher regulatory risk and required the most complete rule evaluation. At a CPU limit of 0.5 cores, the Linux CFS quota mechanism allocated 50 milliseconds of CPU time per 100-millisecond quota period to the validation service's container.

For simple transactions requiring 15–40ms of CPU time, the container completed each request within one or two quota periods and the CFS quota was never exhausted. For complex transactions requiring 350–600ms of CPU time, the container exhausted its 50ms quota multiple times per request: a 500ms CPU task would use 50ms of quota, then be suspended by the CFS for 50ms while the quota period reset, then use another 50ms, then be suspended again, repeating until the full quota was consumed — 500ms of CPU work taking approximately 950ms of wall-clock time, with 450ms of that time spent in the throttled state waiting for the next quota period. The actual throttling overhead depended on the granularity of the CFS quota period and the task's CPU usage pattern; in practice, a 500ms CPU-intensive request on a 0.5-core-limited container took between 850ms and 1,200ms of wall-clock time.

The validation service's P99 latency, measured from Datadog APM spans, was 1.8 seconds. The P50 latency was 85 milliseconds. The 20x gap between median and tail latency had been present for as long as the service had run on Kubernetes — it had never been lower — but it had reached the point where enterprise customers running high-volume reconciliation jobs were experiencing reconciliation delays of 3–4 hours instead of the expected 45 minutes, because the complex transactions in their batches were each taking 1–2 seconds to validate instead of 400ms. The team opened an investigation. They profiled the Go service using pprof: the CPU profiles showed efficient execution with no obvious hotspots; the garbage collector was running at normal frequency. They added distributed tracing: the traces showed no IO wait, no database slowness, no lock contention. The Datadog APM latency heatmap showed a bimodal distribution — a fast cluster at 20–80ms and a slow cluster at 800–1,800ms — that corresponded exactly to the simple/complex transaction split, suggesting the issue was intrinsic to complex request handling rather than a concurrent resource contention problem.

Three weeks into the investigation, an SRE joining from a company with extensive Kubernetes experience asked a single question: "what is your container_cpu_cfs_throttled_seconds_total?" The team queried Prometheus. For the validation service, the throttled_periods / total_periods ratio was 0.67 at the time of the query — 67% of all CPU scheduling periods were throttled. For a service with a 0.5 CPU limit and a 100ms period, this meant the container was spending twice as much time waiting for its quota to reset as it was spending actually running. The complex transactions, which required the most CPU, were the most severely throttled — their wall-clock execution time was dominated by throttle wait rather than CPU execution. The fix was to increase the CPU limit from 0.5 to 2.0 cores. The P99 latency dropped from 1.8 seconds to 310 milliseconds within 15 minutes of the limit change; the throttled_periods ratio dropped to 0.03. The enterprise reconciliation jobs that had been running 3–4 hours completed in 47 minutes. Connect this failure to the observability strategy decision record: the investigation had excellent application-level observability — pprof profiling, distributed tracing, APM latency heatmaps — and none of it was instrumented at the layer where the problem existed; the CFS throttling is a kernel scheduling decision that does not appear in application traces, profiles, or logs; it is visible only through the cAdvisor metrics that the kubelet exports and that a Prometheus scrape collects; an observability strategy that covers application-level signals but not container-runtime scheduling signals has a blind spot at exactly the layer where resource configuration failures manifest.

A 51-person developer tools SaaS company built a code search and navigation platform — fast grep-style search, symbol cross-referencing, and code intelligence features for software repositories. Their backend included an indexing tier — a set of services that parsed source code, built symbol tables, and stored the results in a search index — that ran on a Kubernetes cluster with six nodes, each with 32 GB of memory and 8 CPU cores. The infrastructure team had spent time measuring the indexing service's memory usage and had found that under typical conditions, indexing a standard customer repository used between 200 MB and 800 MB of memory, with an average of approximately 400 MB. To maximize pod density, they had configured the indexing service's memory request at 512 MB — slightly above the average, which they had validated against 30 days of metrics, with a memory limit of 8 GB to accommodate outliers.

The scheduling math had seemed sound: with six nodes at 32 GB each (192 GB total, with approximately 10% reserved for system pods), the cluster could schedule approximately 340 indexing pods at 512 MB each. The team ran 8 indexing pods in production at any given time — a small fraction of theoretical capacity, with plenty of room to scale out via HPA when the queue depth grew. At the 512 MB request, Kubernetes placed 4 indexing pods per node across two of the six nodes. The other four nodes ran the search, API, and storage services. Under typical workload, the eight indexing pods used between 1.6 GB and 6.4 GB of aggregate memory across the two indexing nodes — comfortably within the 32 GB capacity per node.

In the fourth quarter, a new enterprise customer with a large codebase onboarded. Their primary repository contained 3.1 million files across 47 languages, totaling approximately 50 GB of source code. When the customer triggered their first full repository index, the indexing system queued the work and the eight indexing pods began processing partitions of the repository simultaneously. Indexing a partition of the 50 GB repository — the job was divided into eight roughly equal partitions — required approximately 3.2 GB of memory per worker: the indexing service loaded the raw source files, built an in-memory AST for each file, accumulated the symbol table entries, and flushed to the search index in batches; for a partition of this size, the in-memory symbol table alone required 2.4 GB, and the AST buffers added another 0.6–0.8 GB.

All eight indexing pods grew to approximately 3.2 GB of memory usage within 20 minutes of starting the enterprise customer's index. The two indexing nodes, each hosting four pods, were now carrying approximately 12.8 GB of indexing pod memory plus 3–4 GB of system and sidecar processes — roughly 16–17 GB of the 32 GB capacity. The memory usage was not immediately fatal, but it left no headroom for any other process on those nodes. When the garbage collector in the Go indexing process ran during the final flush phase — an operation that temporarily doubled the in-memory data structure size as the old and new representations coexisted — four of the eight pods peaked at 6.4 GB simultaneously. The two indexing nodes, each hosting four pods, ran out of memory at the same time. The Linux OOM killer activated on both nodes simultaneously and killed all four indexing pods on each node — all eight pods were OOMKilled within 14 seconds of each other.

Kubernetes detected the OOMKills and attempted to restart all eight pods. The restart triggered re-scheduling: the two indexing nodes were memory-pressured and were assigned memory-pressure taints by the kubelet, which blocked new pod scheduling onto them. The other four nodes were running the search, API, and storage services at near-full capacity from the enterprise customer's elevated query load following the index start. None of the four healthy nodes had 512 MB of available memory (unscheduled) for the eight indexing pods to reschedule onto — they had available physical memory, but the scheduled capacity (the sum of all other pods' memory requests) had reached the nodes' total capacity from prior scheduling decisions. The eight indexing pods entered Pending state. The enterprise customer's index was abandoned mid-way. Their repository index was left in a partially completed state: 4 of 8 partitions had been successfully indexed and flushed to the search index before the OOMKill, and 4 partitions had been lost; the partial index was loaded but returned incomplete results. The recovery required manually cordoning and draining the two tainted nodes, restarting them to clear the memory pressure taint, and re-queuing the enterprise customer's index — a 4-hour operation during which the customer had access to a partial code search index. Connect this failure to the incident response playbook decision record: the OOMKill event had no pre-defined incident response procedure for the specific failure mode of simultaneous multi-pod OOMKill on an indexing workload; the on-call engineer escalated to the infrastructure team, who escalated to the original cluster designer, who identified the taint mechanism after 35 minutes; the response procedure — cordon the tainted nodes, drain remaining pods, remove the taint, verify node health, re-schedule pods — was reconstructed under incident pressure rather than executed from a documented runbook; a partial index state is a recoverable but data-integrity-affecting condition (search returns wrong results, not no results) that should have a defined customer communication procedure and a specific remediation path in the incident playbook, distinct from a total service outage.

Structural properties set by the Kubernetes resource limits decision

Three structural properties are determined when a team decides — or fails to explicitly decide — what resource requests and limits to configure for their Kubernetes workloads: what the QoS class determines about the blast radius of a single pod's memory failure, what the CPU limits model determines about the latency floor imposed on bursty requests by the CFS quota mechanism, and what the memory request accuracy determines about the over-subscription surface that accumulates as pods are scheduled to nodes based on requests that systematically understate actual peak usage. None of these properties are labeled as decisions at the time the initial Kubernetes manifests are written. The QoS class is set by leaving the resource fields empty because the configuration was deferred. The CPU limits model is set by a policy intended to improve predictability whose CFS throttling implications were not understood. The memory request accuracy is set by a density optimization that used the wrong usage metric as the input.

Property 1: The QoS class and the OOMKill blast radius. The Kubernetes QoS class determines the order in which the Linux OOM killer terminates pods when a node runs out of memory: BestEffort pods (no requests or limits) are killed first, Guaranteed pods (requests equal to limits) are killed last, and Burstable pods (requests below limits) fall in between based on how far their actual usage exceeds their configured request. The blast radius of a single pod's memory failure is the set of pods that are affected when that pod's memory growth forces the OOM killer to intervene. For a Guaranteed pod, the blast radius is bounded by its memory limit: the OOM killer kills the pod when its usage reaches the limit, which happens before the pod's growth can affect other pods' memory availability on the same node. For a BestEffort pod, the blast radius is the entire node: the pod can grow to consume all available memory on the node before the OOM killer acts, at which point the OOM killer may need to kill not just the runaway pod but other pods and system processes that have also consumed memory that is no longer available. The structural difference between a BestEffort cascade (one pod takes down a node) and a Guaranteed containment (one pod is OOMKilled, other pods are unaffected) is entirely a function of whether the runaway pod has a memory limit. A memory limit does not prevent the pod from running out of memory — it does not fix the memory leak that caused the growth — but it constrains the blast radius of the leak to the pod itself, preventing the leak from propagating to the node and to other pods. Connect this property to the service mesh adoption decision record: resource limits and mTLS enforcement share the same structural property — both are containment mechanisms whose protection is binary between "configured and enforced" and "configured but not enforced" or "not configured"; PERMISSIVE mTLS is not zero-trust; BestEffort QoS is not resource management; both provide the operational overhead of the mechanism (sidecar injection, manifest maintenance) without the containment guarantee that motivated the adoption; the containment guarantee requires explicit enforcement, which requires explicit configuration, which requires the explicit decision that the configuration has so far deferred.

Property 2: The CPU limits model and the CFS throttling latency floor. The CPU throttling latency floor is the additional wall-clock time added to requests by the CFS quota mechanism when a container's instantaneous CPU usage exceeds its configured CPU limit within a quota period. The floor is structural: for a container with a 0.5 CPU limit and a 100ms CFS period, the maximum CPU time available in any 100ms window is 50ms; a request requiring more than 50ms of CPU will be throttled regardless of whether the node has spare CPU capacity available; the throttling is imposed by the kernel scheduler, is invisible to the application process (the process is simply suspended), and produces a latency distribution with a long tail at the 95th and 99th percentile that is decoupled from the P50 and median. The latency floor is worst for services whose request CPU usage is multimodal: services with a fast path (low CPU, never throttled) and a slow path (high CPU, heavily throttled) exhibit a P50 that reflects the fast path and a P99 that reflects the slow path with throttling overhead added. This P99/P50 divergence is not visible in average CPU utilization metrics (which reflect the mix of fast and slow requests) and is not visible in application-level traces (which measure wall-clock time but do not distinguish throttle wait from actual processing). The only signal is the CFS throttling metric from the container runtime. The correct model for latency-sensitive services with bursty CPU usage is to set no CPU limit or to set the CPU limit well above the measured P99 CPU usage per request, accepting that the service can burst to available node CPU capacity at the cost of potentially consuming a neighbor's spare capacity. For services where predictable CPU consumption is more important than tail latency (batch jobs, background workers), CPU limits equal to requests are appropriate. The decision of which model applies to which service is a service classification decision that must be made explicitly at service onboarding, not implied by a uniform policy applied to all services without regard to their CPU usage distribution. Connect this property to the alerting threshold decision record: the CFS throttling ratio (throttled_periods / total_periods) is a CPU resource signal that has a direct latency interpretation — at 10% throttling, tail latency is elevated; at 50% throttling, tail latency is severe; the alert threshold for this metric must be set as an absolute value (alert at 10% throttling ratio), not derived from historical baselines, because historical baselines will be calibrated to the throttled state; an alert that fires when throttling ratio rises above historical average will fire too late, after the performance has already degraded to its throttled floor.

Property 3: The memory request accuracy and the over-subscription cascade surface. The memory over-subscription cascade surface is the set of nodes where the sum of scheduled pods' actual peak memory usage exceeds the node's physical capacity when those peaks coincide. The surface exists whenever memory requests are set below the actual peak usage of the scheduled pods, because the Kubernetes scheduler places pods based on requests — not actual usage — and the gap between requested and actual peak creates capacity that the scheduler believes is available but that the node cannot actually serve. The cascade trigger is any event that drives correlated peak memory usage across pods that normally peak at different times: a large customer, a scheduled batch job, a seasonal traffic spike, a correlated data export — any workload that drives all pods toward their peak simultaneously rather than distributing peaks across time. The cascade surface is proportional to the aggregate under-declaration: eight pods each under-declaring by 6x (512 MB request vs 3.2 GB peak) create a 20.8 GB aggregate over-subscription on two nodes that each have 32 GB of physical memory — the nodes appear to have 47% utilization from the scheduler's perspective while actually having 0% headroom at peak. The correct model uses the 95th-percentile actual peak usage as the memory request, with the limit set at 1.5–2x the request. Setting requests at the 95th percentile does not eliminate the over-subscription risk for workloads with extreme outliers (larger than 5th percentile), but it eliminates the systematic over-subscription that accumulates when requests are set at average or median usage. For workloads with extreme outlier potential (services whose memory usage is proportional to a customer-supplied input size, like the indexing service in story 3), the request should be set at the P99 usage for known typical inputs, and the limit should be set based on the maximum expected input size — not on the distribution of typical inputs. Connect this property to the capacity planning decision record: memory request accuracy is a capacity planning input, not a cluster configuration parameter; the cluster's effective capacity is the minimum of its physical capacity and the aggregate committed capacity (the sum of all requests); if requests are set below actual usage, the cluster's scheduler sees a higher effective capacity than the cluster can actually serve, and the cluster appears to have headroom that does not exist; capacity planning that uses scheduled requests as the capacity metric will conclude the cluster is 45% utilized when it is 95% utilized at peak, producing a safety margin that is an artifact of under-declaration rather than a genuine buffer.

The Kubernetes resource limits ADR: five sections

Section 1: QoS class policy and service classification. Begin the Kubernetes resource limits decision record by specifying the QoS class policy for each service type and the criteria for classification. Three service types require distinct treatment. Stateful services and data-tier components (databases, caches, message brokers, stateful workers where OOMKill causes data loss or requires manual recovery): configure Guaranteed QoS — memory request equal to memory limit, CPU request equal to CPU limit; the Guaranteed class ensures these pods are evicted last under memory pressure and their resource consumption is completely predictable; set the memory limit at the P99 observed usage plus 25% buffer; set the CPU limit at a value well above the P99 CPU usage to avoid CFS throttling on these typically non-bursty services. Latency-sensitive application services (API handlers, synchronous request processors, user-facing endpoints where tail latency is a product quality metric): configure Burstable QoS with no CPU limit or with a high CPU limit (at least 3x the P95 CPU usage per request); omitting the CPU limit allows the service to burst to available node CPU capacity on high-complexity requests without throttling; set the memory request at P95 observed usage and the memory limit at 2x the request; the Burstable class with a high memory limit allows these services to absorb request complexity spikes without OOMKill while preventing unbounded growth. Background processing services (batch workers, indexing services, export generators, async job processors): configure Burstable QoS with CPU limits appropriate to their acceptable throughput target, and memory requests set at the P99 usage for the expected maximum workload unit; for services whose memory usage scales with input size, set the request and limit based on the maximum input size the service should handle, not on typical input sizes. Specify a LimitRange resource in each namespace that sets default requests and limits for containers that do not specify them — this prevents new services from running as BestEffort by accident when developers don't set resource fields. Connect to the container orchestration decision record: the QoS class policy should be specified in the same decision record as the cluster topology — the number of nodes, their instance types, and the pod density target; QoS class determines which workloads the cluster can co-schedule safely, and the cluster topology determines the blast radius containment at the node level; a single-node cluster with BestEffort pods has no blast radius containment (any pod failure can take the cluster down); a multi-node cluster with Guaranteed stateful services and co-scheduled BestEffort batch jobs has defined containment (the BestEffort batch job can take the node, not the cluster).

Section 2: CPU request and limit configuration model. Specify the CPU configuration model for each service classification, the CFS quota period to use (or disable), and the monitoring approach for detecting throttling. For latency-sensitive services: set the CPU request at the P50 CPU usage per request times the target requests per second the pod should handle; this gives the scheduler a realistic view of the CPU the pod will consume under normal load; set no CPU limit, or set the CPU limit high enough (4x the CPU request) that CFS throttling does not occur at the P99 request complexity; justify the choice between no-limit and high-limit: no-limit maximizes burst capacity but allows runaway CPU consumption on a node; high-limit contains runaway CPU but imposes throttling at the tail; for most latency-sensitive services, no-limit is the correct choice because node-level CPU over-subscription is far less catastrophic than node-level memory over-subscription (CPU contention degrades performance but does not cause crashes). For CPU-intensive batch services: set the CPU limit at a value that achieves the acceptable throughput within the scheduled time window; a batch job with a 4-hour completion requirement needs enough CPU to process its workload in 4 hours; calculate the required average CPU from the total CPU-seconds of work divided by the time budget; set the limit at 1.2x this value. Specify the CFS period: the default 100ms CFS period is not configurable per-container in most Kubernetes versions, but the scheduler's CPU accounting operates on this period; for services with bursty CPU usage, understand that a 0.5 CPU limit means 50ms of CPU per 100ms period — any request requiring more than 50ms of continuous CPU will be throttled; if the CFS throttling metric shows persistent throttling above 5% on a latency-sensitive service, the first response is to increase the CPU limit, not to tune the application. Specify the throttling alert: alert on container_cpu_cfs_throttled_periods_total / container_cpu_cfs_periods_total > 0.10 for any latency-sensitive service; this is a 10% throttling ratio threshold that provides early warning before tail latency degradation is visible to users. Connect to the observability strategy decision record: add container-runtime resource metrics (container_cpu_cfs_throttled_periods_total, container_memory_working_set_bytes, container_memory_oom_kills_total) to the observability strategy's signal inventory alongside application-level traces and logs; these metrics are not surfaced by default in most APM tools and require explicit Prometheus scrape configuration; a service health definition that does not include container-runtime signals will produce correct application-layer alerts but will miss the class of failure (CFS throttling, memory pressure, OOMKill) that originates at the container scheduling layer.

Section 3: Memory request and limit policy. Specify the memory request policy and the measurement methodology for determining request values. The request should be set at the 95th-percentile memory usage measured over a representative traffic period (at least 7 days, including any weekly periodic patterns such as weekly report generation or scheduled batch jobs that drive above-average memory usage). Do not use average or median memory usage as the request value — average usage produces systematic over-subscription because pods regularly use more than their request, which misleads the scheduler into placing more pods on a node than the node can safely host at peak. For services with input-dependent memory usage (services that process customer-uploaded data, services that index customer repositories, services whose memory scales with query complexity): measure peak memory usage at the maximum expected input size, not at the average input size; set the request at the peak for a representative large customer; accept that this reduces node density relative to average-input-sized requests, because the alternative is a crash when the large customer's input triggers peak usage across multiple pods simultaneously. Specify the memory limit for each service type: for latency-sensitive services, set the memory limit at 2x the request; for batch services, set the limit at the maximum expected usage for the largest job the service will run; for stateful services, set the limit equal to the request (Guaranteed QoS). Specify the over-subscription budget: define the maximum ratio of total requests to total node capacity that is acceptable (recommended: no more than 80% of node capacity in aggregate requests, leaving 20% for burst headroom); this budget is a scheduling constraint that the cluster's autoscaler should enforce by adding nodes before the scheduled-capacity-to-physical-capacity ratio exceeds the budget. Connect to the capacity planning decision record: memory requests are the interface between the service configuration and the capacity model; a capacity model that uses physical memory utilization (actual usage) as its primary signal will read 45% utilization and see headroom; a capacity model that uses scheduled capacity (sum of requests) as its primary signal will see 93% scheduled capacity and trigger node provisioning; both models are necessary — scheduled capacity tracks scheduling safety, physical utilization tracks actual resource consumption; the gap between scheduled capacity and physical utilization is the measurement of over-declaration (requests above actual usage) or under-declaration (requests below actual usage); over-declaration wastes capacity; under-declaration creates the over-subscription cascade surface.

Section 4: Workload-specific configuration for variable-memory workloads. Specify the configuration model for services whose memory usage varies with the input they process — services where a customer-supplied upload, a large query, or an unusually sized dataset can drive memory usage far above the typical workload's usage. Three mechanisms address variable-memory workloads. First, input size gating: restrict the maximum input size the service will accept to a value that bounds the maximum memory usage below the configured limit; if the indexing service has a 4 GB memory limit, accept only repositories whose size is bounded below the threshold that produces 4 GB usage; reject oversized inputs at the API layer with a clear error message rather than allowing the service to be OOMKilled. Second, job isolation: for workloads where input size cannot be bounded (customer-supplied data must be fully processed), run large inputs as isolated Kubernetes Jobs on dedicated nodes rather than on the shared service nodes; configure Taints on the dedicated nodes that repel all workloads except the large-input jobs, and configure Tolerations on the large-input Jobs to allow them to land on the dedicated nodes; this prevents a large-input job from competing for memory with production service pods on shared nodes. Third, Vertical Pod Autoscaler (VPA): configure VPA in UpdateMode: Off (recommendation mode only) to collect memory usage measurements from running pods; the VPA recommendation provides a data-driven starting point for request and limit values based on observed usage; periodically review the VPA recommendations and update the request values if the recommendation diverges by more than 20% from the configured value; do not use VPA in automatic enforcement mode on production services without extensive testing, because automatic request updates can trigger pod restarts at unexpected times. Connect to the multi-tenant data isolation decision record: a 50 GB customer repository as a memory usage trigger is a tenant-driven resource isolation problem as well as a Kubernetes configuration problem; the resource limits decision determines the containment at the node level, but the decision of whether to run large-tenant workloads in the same scheduling context as other tenants' workloads is a multi-tenant isolation decision; running all tenants' indexing jobs on the same node pool with shared QoS creates a noisy-neighbor surface where one large tenant's workload affects all other tenants' indexing performance; dedicated node pools per tenant tier (enterprise tenants with large repositories on dedicated nodes, standard tenants on shared nodes) is the multi-tenant isolation mechanism that the resource limits decision complements.

Section 5: Monitoring, alerting, and LimitRange enforcement. Specify the monitoring requirements for resource limit health, the alert thresholds for each signal, and the namespace-level enforcement mechanisms that prevent services from running with misconfigured resources. The monitoring requirements: container_memory_working_set_bytes / container_spec_memory_limit_bytes — alert at 80% of limit per container, indicating approaching OOMKill risk; container_cpu_cfs_throttled_periods_total / container_cpu_cfs_periods_total — alert at 10% throttling ratio per container, indicating CFS throttling is affecting tail latency; kube_pod_container_resource_requests with requests = 0 — alert immediately, indicating a container is running as BestEffort without intentional classification; kube_pod_container_status_restarts_total with reason = OOMKilled — alert on any OOMKill, regardless of severity, with the pod name and namespace to enable root cause investigation; node_memory_available_bytes — alert when node available memory drops below 2 GB, providing earlier warning than the per-pod memory alerts. Specify the LimitRange: configure a LimitRange in each production namespace that sets default requests and limits for containers that do not specify them, preventing new services from running as BestEffort by default. A reasonable default: CPU request 100m, CPU limit (no limit is preferable for latency-sensitive services — set the LimitRange to set a high default limit of 2 CPUs), memory request 256 MiB, memory limit 1 GiB. The LimitRange default is a starting point; services that have measured their actual usage should override these defaults with accurate values. Specify the resource accuracy review cadence: quarterly, run a review that compares each service's configured requests against its actual P95 usage from the previous 90 days; services where the configured request deviates from actual P95 usage by more than 50% (either direction) should have their requests updated; the review is a calendar item with an assigned owner, not an ad hoc response to incidents. Connect to the incident response playbook decision record: add OOMKill and node memory pressure to the incident playbook as Tier 2 incidents with a defined response procedure: (1) check kube_node_status_condition for the memory-pressure taint; (2) run kubectl top nodes to identify which node is memory-pressured; (3) run kubectl top pods -A --sort-by=memory on the pressured node to identify the highest-memory consumers; (4) check if any pod is approaching its limit; (5) if the OOMKill cycle is ongoing, cordon the node to prevent new scheduling, drain remaining pods, and remove the taint after the memory-leaking pod has been identified and its limit has been enforced; the playbook makes the difference between an 80-minute cascade and a 12-minute containment.

FAQ

What Kubernetes QoS class should each type of service use?

Use Guaranteed QoS (requests equal to limits on every container) for stateful services and data-tier components — databases, caches, stateful workers — where OOMKill causes data loss or requires manual recovery; Guaranteed pods are evicted last under memory pressure, protecting the components most expensive to recover. Use Burstable QoS (at least one container has a request or limit, but they are not equal) for latency-sensitive API services: set the memory request at P95 observed usage and the memory limit at 2x the request; omit the CPU limit or set it well above P95 CPU usage per request to avoid CFS throttling on bursty workloads. Never run continuously-running production services as BestEffort (no requests or limits) — BestEffort pods have the highest oom_score and are the first to be killed under memory pressure; they are appropriate only for batch jobs and workloads that can tolerate eviction and restart. Configure a LimitRange in each production namespace that sets safe default requests and limits so that services without explicit configuration run as Burstable rather than BestEffort. For CPU specifically: avoid setting CPU limits equal to CPU requests on latency-sensitive services; this creates Guaranteed CPU QoS but imposes CFS throttling on any request that requires more than one CFS quota period's worth of CPU, which is any request above the CPU limit value times 100ms; for bursty services, no CPU limit is preferable to an equal-to-request CPU limit.

How do you detect and resolve CFS CPU throttling that is causing tail latency problems?

Query container_cpu_cfs_throttled_periods_total / container_cpu_cfs_periods_total from Prometheus (cAdvisor metrics, available from most Kubernetes monitoring setups). This ratio is the fraction of CFS quota periods during which the container was throttled. A ratio above 5% indicates throttling is occurring; above 25% indicates measurable tail latency impact; above 50% means the container is spending more time waiting for its quota to reset than actually running. Alert at 10% as an early warning. The throttling is invisible to application-level profiling and tracing — pprof, APM spans, and distributed traces measure elapsed wall-clock time but cannot distinguish CPU execution time from CFS throttle wait; the CFS metrics are the only reliable signal. To resolve: increase the CPU limit, or remove it. Increasing the CPU limit (e.g., from 0.5 to 2.0 cores) raises the quota ceiling and allows bursty requests to consume more CPU within a single quota period before being throttled; this is the correct first response for services with a defined burst maximum. Removing the CPU limit entirely allows the container to burst to available node CPU capacity without any throttling; this is appropriate for latency-sensitive services where tail latency matters more than guaranteed CPU isolation from neighbors. After increasing or removing the CPU limit, verify the throttling ratio drops below 5% and confirm the improvement in tail latency via the P99 latency metric. Update the CPU request if the service's measured average CPU usage has changed — the request should reflect typical usage, not the limit.

How do you set memory requests to prevent node over-subscription without wasting capacity?

Set memory requests at the 95th-percentile actual memory usage measured from production, not at average or median. Average usage creates systematic over-subscription: if half of running instances regularly exceed their request, the scheduler places more pods on a node than it can safely host at peak. For services with input-dependent memory usage (indexers, processors, exporters), measure peak usage at the maximum expected input size and set the request based on that peak, not on typical inputs. The trade-off: requests set at P95 usage reduce pod density per node (fewer pods fit at the scheduled capacity per node) but eliminate the over-subscription cascade triggered by simultaneous peak usage events. The density trade-off is worth it: the cost of lower density is slightly higher infrastructure cost; the cost of over-subscription is cascading OOMKills, node pressure taints, and multi-hour recovery. Set memory limits at 1.5–2x the request to allow headroom for usage above P95 without allowing unbounded growth. Monitor container_memory_working_set_bytes for each pod and alert at 80% of limit. Use VPA in recommendation mode to collect usage measurements and flag pods whose configured requests diverge by more than 50% from measured P95 usage; the VPA recommendation is a quarterly review input, not an automatic update trigger. For nodes, set a scheduling over-subscription budget: do not schedule requests that total more than 80% of node physical memory, leaving 20% for burst headroom; configure the cluster autoscaler to add nodes before the 80% scheduled-to-physical ratio is reached.

How should Horizontal Pod Autoscaler scaling targets account for resource request accuracy?

Set HPA target utilization based on the relationship between the request and the service's actual capacity: if the request is set at P95 usage (as recommended), an HPA target of 70% means "scale out when average usage exceeds 70% of P95 usage" — which triggers scale-out before pods hit their P95 usage ceiling, providing headroom for the 5% of requests above P95. If the request is set too low (at average usage), the HPA will see 100% utilization at normal load and scale out constantly; if set too high (at P99 usage), the HPA will see low utilization and scale in aggressively until a traffic spike reveals the mismatch. The correct configuration sequence: (1) measure P95 actual usage; (2) set request at P95; (3) set HPA target at 70% of request; (4) set minReplicas to the count needed to handle traffic if one pod is unavailable (to prevent scale-in to a single pod that cannot absorb a restart); (5) set maxReplicas based on the maximum expected traffic spike and the cluster's ability to accommodate the additional pods. For memory-based HPA: memory usage often does not scale proportionally with request rate (many services retain allocated memory even when request rate decreases), making memory-based HPA unreliable for scale-in decisions; prefer CPU-based HPA for most web services; use memory-based HPA only for services whose memory usage reliably scales with request count and reliably releases memory when requests decrease, such as in-memory batch processors.

Further reading

  • Capacity planning decision record — the cluster capacity model that determines how scheduled capacity (requests) and physical utilization are measured and monitored as separate signals; a capacity plan that uses only physical utilization misses the over-subscription surface that accumulates from under-declared requests; a plan that uses only scheduled capacity misses the actual memory pressure that manifests when pods simultaneously reach their peaks; the two signals together define the safe operating range for a Kubernetes cluster and the threshold at which new nodes must be added.
  • Observability strategy decision record — the signal inventory that must include container-runtime metrics (CFS throttling ratios, memory working set relative to limit, OOMKill events) alongside application-level metrics and traces; application traces are blind to CFS throttling (the process is suspended, not executing, so no trace span is emitted for the throttle wait); the container runtime metrics layer is the only source of signal for the class of failure that originates in the Linux kernel scheduler; an observability strategy that does not explicitly include cAdvisor metrics in its signal inventory will diagnose resource configuration failures as application bugs.
  • Incident response playbook decision record — the tier classification and response procedure for OOMKill and node memory pressure events; an OOMKill cascade (multiple pods OOMKilled simultaneously, node enters NotReady) is structurally different from a single-pod OOMKill (one pod restarted, service briefly degraded) and requires a different response procedure; the playbook must specify the node cordon-drain-taint-removal sequence for cascade events and distinguish it from the simpler restart-and-monitor response for isolated OOMKills.
  • Service mesh adoption decision record — the structural parallel between mTLS enforcement mode and QoS class: both are containment mechanisms that are either configured and enforced (STRICT mTLS / Guaranteed QoS) or configured but not enforced (PERMISSIVE mTLS / BestEffort QoS); both incur the operational overhead of the mechanism without delivering the containment guarantee that motivated adoption; both require explicit configuration decisions to move from the permissive default to the enforced state, and both decisions are deferred indefinitely when the motivation is present but the urgency is not.
  • Alerting threshold decision record — the alert calibration model for CFS throttling and memory pressure signals; CFS throttling alerts should be set at absolute thresholds (10% throttling ratio = warning, 25% = page) rather than relative-to-baseline thresholds, because a baseline calibrated against a throttled state will normalize the performance degradation; memory working set alerts should be set at fixed percentages of the configured limit (80% = warning, 90% = page) to provide consistent lead time before OOMKill regardless of the absolute limit value.
  • Open-source extractor — find the Kubernetes resource configuration decisions buried in your AI chat history: the session where resource fields were left unset because measuring actual usage was deferred, the conversation where CPU limits were set equal to CPU requests for "predictability" without discussing CFS quota periods, and the planning discussion where memory requests were set at average usage to maximize density without accounting for simultaneous peak usage events — each of these is a recoverable decision record that explains the structural gap the next resource exhaustion incident will exploit.