Kubernetes Workload Reliability
44 incident problems about Kubernetes Workload Reliability. Start with the reviewed ones.
Read first
When the ConfigMap changed but the Pod keeps using the old valuesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when values don't refresh after a ConfigMap update because of envFrom, subPath, volume projection, or rollout trigger differences.Kubernetes3 min readWhen CoreDNS is Running but only DNS lookups failAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the CoreDNS Pod looks healthy but failures occur in the kube-dns path, upstream, node-local-dns, or policy.Kubernetes3 min readWhat to check first when CrashLoopBackOff appearsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a Pod restarts repeatedly and the real cause is left in the previous logs and events.Kubernetes3 min read
Recommended problems
Reviewed problems first, then problems with detailed scenarios.
CrashLoopBackOff: pods keep restarting after a new releaseCrashLoopBackOff: pods keep restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read the previous container's log to find why a pod keeps restarting. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current liv...ReviewedKubernetesBeginner3 minFreeExit code 137 (OOMKilled): a container restarts under loadExit code 137 (OOMKilled): a container restarts is a hands-on troubleshooting drill. Learn what OOMKilled and exit code 137 mean. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는 CrashLoop...ReviewedKubernetesBeginner3 minFreeOOMKilled: a container keeps restarting after hitting its memory limitOOMKilled: a container keeps restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read OOMKilled and exit code 137, then fix memory limits and the workload that exceeds them. kubernetes-workload-reliability needs to be checked by narrowing scope, recent cha...ReviewedKubernetesBeginner15 minFree
All problems (44)
CrashLoopBackOff: pods keep restarting after a new releaseCrashLoopBackOff: pods keep restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read the previous container's log to find why a pod keeps restarting. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current liv...ReviewedKubernetesBeginner3 minFreeExit code 137 (OOMKilled): a container restarts under loadExit code 137 (OOMKilled): a container restarts is a hands-on troubleshooting drill. Learn what OOMKilled and exit code 137 mean. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는 CrashLoop...ReviewedKubernetesBeginner3 minFreeOOMKilled: a container keeps restarting after hitting its memory limitOOMKilled: a container keeps restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read OOMKilled and exit code 137, then fix memory limits and the workload that exceeds them. kubernetes-workload-reliability needs to be checked by narrowing scope, recent cha...ReviewedKubernetesBeginner15 minFreeK8S-106CronJob timezone is configured but the controller resumes a suspended object on the next stale schedule windowOperators unsuspend a job expecting the new timezone to apply cleanly, yet the controller immediately runs an old missed schedule slot.KubernetesIntermediate15 minFreeK8S-110Indexed Job completions succeed except one shardThe workload mostly finishes, but one index loops forever because the controller and configuration naming scheme disagree on shard-specific inputs.KubernetesIntermediate15 minFreeK8S-120Node feature label disappears during upgrade and the GPU DaemonSet silently unschedules from every new workerThe nodes join successfully, yet the supporting workload never follows because the upgrade process changed or dropped the advertised feature label.KubernetesIntermediate16 minFreeK8S-122Pod anti-affinity is only preferred, so after a zone loss the replacement replicas pile onto one node and the next failure takes out the whole serviceAvailability degrades quietly because the scheduler still honors availability under pressure, just not in the way the team assumed.KubernetesIntermediate17 minFreeK8S-159A canary Service routes to the right pods, but topology spread on the new ReplicaSet sends all canary capacity to one zone and synthetic tests misjudge global readinessTraffic experiments are statistically misleading because one zone now dominates the canary population.KubernetesAdvanced17 minProK8S-112HorizontalPodAutoscaler sees the custom metric intermittentlyAutoscaling seems random because the adapter is reachable, yet TLS validation fails sporadically between control plane components.KubernetesAdvanced18 minProK8S-101StatefulSet ordinal restarts collide with the PodDisruptionBudget and quorum never reforms after node maintenanceOne replica keeps waiting for another to recover, but eviction protections prevent the exact restart sequence the quorum algorithm needs.KubernetesAdvanced19 minProK8S-369A drained node returns to service (Resource Exhaustion)A drained node returns to service (Resource Exhaustion) focuses on kubernetes-workload-reliability and asks the reader to isolate Resource Exhaustion. 실무에서는 resource-exhaustion 증상만 보고 Pod 하나에 매달리지 말고 이벤트, 이전 로그, Service/Endpoint, 최근 배포 변경을 한 번에 묶어 보는 편이 오진을 줄입니다.KubernetesAdvanced17 minProK8S-157A pod restart budget is healthy, but the cluster autoscaler removes the only node with local PV affinity and the replacement pod cannot rescheduleCapacity remains, yet storage locality still pins recovery to a node that no longer exists.KubernetesAdvanced17 minProK8S-148An HPA based on external queue depth scales up correctly, but scale-down never happensThe autoscaler is healthy, yet one observability component keeps a historical answer longer than expected.KubernetesAdvanced17 minProK8S-118PodDisruptionBudget healthy count looks satisfied but drain still blocks on an unready terminating podMaintenance cannot complete because a pod in termination keeps being counted in one state and excluded in another, confusing the disruption math.KubernetesAdvanced17 minProK8S-121A DaemonSet surge update doubles the hostPort bind and the new pods never become Ready on half the nodesThe rollout strategy looks safer on paper, but the host-level port exclusivity means old and new pods cannot overlap.KubernetesAdvanced18 minProK8S-216A local storage workload looks healthy until scale-down removes the only node that satisfies its locality during a failover rehearsalCompute capacity remains while schedulability disappears because the data stayed behind. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced18 minProK8S-141A node drain respects the PodDisruptionBudget, but a custom controller recreates helper pods instantly and the drain never convergesKubernetes is behaving correctly, yet another controller keeps replenishing the very pods maintenance is trying to evict.KubernetesAdvanced18 minProK8S-160A node image upgrade enables cgroup v2, and one Java workload starts misreporting heap limits so the HPA scales on the wrong utilization basisThe app still runs, but runtime memory accounting changed with the node image and distorts autoscaling decisions.KubernetesAdvanced18 minProK8S-345A topology spread rule looks balanced while one tainted node pool is the only place the emergency fallback pods are actually allowed to run during a staged decommissionDistribution appears fair until a failure forces the scheduler into the only domain with compatible taints. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.KubernetesAdvanced18 minProK8S-204An autoscaler trusts a queue-depth metric that remains stale long after backlog is gone during a failover rehearsalScaling keeps behaving as if yesterday's pressure still exists because freshness is not part of the metric contract. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced18 minPro