Anthony Linsday
[ BACK TO HOME ]   STATUS: ONLINE
< BACK TO HOME
ENGINEERING DISPATCH // POST 0x42

Kubernetes in Production: Battle-Tested Patterns

6 MIN READ
AUTHOR: Anthony Linsday DATE: Jan 28, 2026 CATEGORY: Container Orchestration & EKS

// 01. The Trap of Rigid CPU Limits

One of the most insidious causes of high P99 latency in containerized applications is misconfigured CFS (Completely Fair Scheduler) CPU quotas. When you declare resources.limits.cpu: "500m", the Linux kernel enforces hard CPU throttling per 100ms CFS quota window. Even if the host node has 60% idle capacity, your Node.js or Java threads will freeze mid-request.

Rule of thumb: In production web services, set conservative requests.cpu to ensure node scheduler bin-packing, and consider omitting limits.cpu (or setting it generous) while relying on memory limits to prevent out-of-memory cascades.

// 02. Graceful Shutdown & PreStop Hooks

When a pod is terminated during a rolling update or node drain, Kubernetes removes the pod from endpoints and sends SIGTERM concurrently. Because kube-proxy iptables propagation has a delay of 1 to 3 seconds, incoming client traffic will still route to your dying pod.

lifecycle:
  preStop:
    exec:
      command: ["/bin/sh", "-c", "sleep 5"]
# Allows ingress controller and kube-proxy iptables to drain connections
# before the application server initiates its graceful close sequence.

// 03. PodDisruptionBudgets (PDB) are Mandatory

Node upgrades, security patch cycles, and Spot instance terminations will wreak havoc on your clusters without PDBs. For any multi-replica service:

  • Always configure minAvailable: 1 or maxUnavailable: 25%.
  • Enforce pod anti-affinity so replicas span distinct Availability Zones (AZs).
  • Use Karpenter or Cluster Autoscaler with consolidation enabled to optimize cloud compute expenditure.