Kubernetes in Production: Battle-Tested Patterns
// 01. The Trap of Rigid CPU Limits
One of the most insidious causes of high P99 latency in containerized applications is misconfigured CFS (Completely Fair Scheduler) CPU quotas. When you declare resources.limits.cpu: "500m", the Linux kernel enforces hard CPU throttling per 100ms CFS quota window. Even if the host node has 60% idle capacity, your Node.js or Java threads will freeze mid-request.
Rule of thumb: In production web services, set conservative requests.cpu to ensure node scheduler bin-packing, and consider omitting limits.cpu (or setting it generous) while relying on memory limits to prevent out-of-memory cascades.
// 02. Graceful Shutdown & PreStop Hooks
When a pod is terminated during a rolling update or node drain, Kubernetes removes the pod from endpoints and sends SIGTERM concurrently. Because kube-proxy iptables propagation has a delay of 1 to 3 seconds, incoming client traffic will still route to your dying pod.
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"]
# Allows ingress controller and kube-proxy iptables to drain connections
# before the application server initiates its graceful close sequence. // 03. PodDisruptionBudgets (PDB) are Mandatory
Node upgrades, security patch cycles, and Spot instance terminations will wreak havoc on your clusters without PDBs. For any multi-replica service:
- Always configure
minAvailable: 1ormaxUnavailable: 25%. - Enforce pod anti-affinity so replicas span distinct Availability Zones (AZs).
- Use Karpenter or Cluster Autoscaler with consolidation enabled to optimize cloud compute expenditure.