Kubernetes
586 incident problems in Kubernetes environments.
먼저 읽을 가이드
추천 문제
All problems (586)
K8S-1336A pod resolves services on old nodes and times out on fresh nodesA DNS service migration completes and later only pods on newly added nodes see timeouts or stale resolution results.KubernetesIntermediate13 minProK8S-1324A PodDisruptionBudget blocks every node drainA PodDisruptionBudget blocks every node drain focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Terminating is not the same as absent from the budget calculation.KubernetesIntermediate13 minProK8S-1302A postStart hook was used as a startup gate and traffic arrives too earlyA team hides bootstrap steps inside a lifecycle hook and later sees early requests fail despite no pod crash.KubernetesIntermediate13 minProK8S-1358A service mesh sidecar rollout appears safe and readiness flapsA sidecar injection rollout is enabled and later pods begin failing readiness despite healthy application logs.KubernetesIntermediate13 minProK8S-1332A ServiceMonitor appears valid and Prometheus still scrapes nothingA ServiceMonitor appears valid and Prometheus still scrapes nothing focuses on runtime-configuration and asks the reader to isolate the key signal in Kubernetes. Prometheus can miss perfectly valid monitor objects if they moved outside the...KubernetesIntermediate13 minProK8S-1354A cert-manager renewal job stays Pending (webhook-networkpolicy-still-whitelisted-old-apiserver-cidr)A cert-manager renewal job stays Pending (webhook-networkpolicy-still... focuses on Identity And Access and asks the reader to isolate the key signal in Kubernetes. Admission and webhook traffic often comes from a control-plane CIDR...KubernetesIntermediate14 minProK8S-1326A cert-manager renewal succeeds and the new certificate never reaches the appA certificate rotates successfully in cluster state and the application continues presenting the old material until a restart.KubernetesIntermediate14 minProK8S-1338A Cilium upgrade keeps policy enforcement and breaks one NodePort pathA Cilium upgrade keeps policy enforcement and breaks one NodePort path focuses on network-segmentation and asks the reader to isolate the key signal in Kubernetes. Partial CNI rollouts can create path-specific behavior differences th...KubernetesIntermediate14 minProK8S-1313A CronJob fires twice (cronjob-duplicate-run-after-eviction)A CronJob fires twice (the key signal) focuses on incident-response and asks the reader to isolate the key signal in Kubernetes. Cluster scheduling guarantees do not replace application-level idempotency for CronJobs.KubernetesIntermediate14 minProK8S-1334A node drain hangs even though replicas are healthyA maintenance window begins and a single namespace blocks node drains with admission errors unrelated to application health.KubernetesIntermediate14 minProK8S-1329A Pod logs DNS timeouts only on fresh nodesA Pod logs DNS timeouts only on fresh nodes focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Node-specific DNS failures often point to bootstrap order rather than to cluster-wide resolver health.KubernetesIntermediate14 minProK8S-1344A StatefulSet pod comes up and fails readinessA StatefulSet pod comes up and fails readiness focuses on runtime-configuration and asks the reader to isolate the key signal in Kubernetes. Stateful recovery can fail on identity artifacts even when the raw data volume is fine.KubernetesIntermediate14 minProK8S-1322A StatefulSet pod stays Pending after force deletionAn operator force deletes a stuck StatefulSet pod during an incident and later the replacement remains Pending with attach-related events.KubernetesIntermediate14 minProK8S-1308An ingress controller never reloads one new configA platform team separates controller and configuration namespaces and later one class of ingress settings silently stops applying.KubernetesIntermediate14 minProK8S-1301An initContainer loops on DNS for a dependency that already existsA rollout creates both a dependency service and a consuming workload, and only the first few pods never recover from DNS failures.KubernetesIntermediate14 minProK8S-1620A PVC expansion succeeds and one pod still reports no spaceOne pod still reports no space after PVC expansion.KubernetesIntermediate10 minProK8S-1630NodeLocal DNS is healthy and one node still resolves stale namesOne node still resolves stale names after a CoreDNS upstream change.KubernetesIntermediate10 minProK8S-1700A Cilium policy is correct and one DNS path still failsA DNS path still fails from one workload even after the Cilium policy was corrected.KubernetesIntermediate11 minProK8S-1790A mesh route is healthy and one client still failsOne client still fails even though the mesh route is healthy.KubernetesIntermediate11 minProK8S-1810A mesh route is healthy and one client still failsOne client still fails even though the mesh route is healthy.KubernetesIntermediate11 minPro