Vendor586 problems· 4 reviewed

Kubernetes

586 incident problems in Kubernetes environments.

All problems (586)

K8S-1423A cluster autoscaler adds nodes and pending pods still stay unscheduledA cluster autoscaler adds nodes and pending pods still stay unscheduled focuses on scheduler-behavior and asks the reader to isolate the key signal in AWS. Capacity can look available on paper while scheduler caches and...KubernetesAdvanced14 minProK8S-1413A cluster autoscaler scale-up looks correct and pending pods remain...A cluster autoscaler scale-up looks correct and pending pods remain... focuses on scheduler-behavior and asks the reader to isolate the key signal in AWS. Autoscaler failures can come from stale node template metadata rather than from a...KubernetesAdvanced14 minProK8S-1433A scale-up event adds nodes and pending pods remain unscheduledA scale-up event adds nodes and pending pods remain unscheduled focuses on scheduler-behavior and asks the reader to isolate the key signal in AWS. New nodes can be present and still unusable when simulation metadata lags behind the act...KubernetesAdvanced14 minProK8S-1443A scale-up says pending pods will fit and they stay unscheduledA node pool migration completes and target workloads remain pending.KubernetesAdvanced14 minProK8S-1349A Cilium Hubble flow view shows allowed traffic and the app still failsOperators investigate a connectivity issue with Hubble and later discover the visible pod flows were never the part of the path doing the drop.KubernetesAdvanced15 minProK8S-1361A Cilium rollout leaves one namespace unreachableA platform team standardizes labels and later one namespace loses east-west traffic only on a subset of nodes.KubernetesAdvanced15 minProK8S-1371A Cilium upgrade leaves one namespace unreachableA label cleanup is rolled out and later one namespace loses traffic only on a subset of nodes.KubernetesAdvanced15 minProK8S-1327A DaemonSet appears healthy and one node family never gets the agentA fleet adds a new node pool image and later one DaemonSet quietly skips those nodes while remaining healthy elsewhere.KubernetesAdvanced15 minProK8S-1392A Gatekeeper policy looks unchanged and one namespace starts failing admissionA Gatekeeper policy looks unchanged and one namespace starts failing admission focuses on Identity And Access and asks the reader to isolate the key signal in Kubernetes. Admission drift can hide in CRD conversion...KubernetesAdvanced15 minProK8S-1402A Gatekeeper policy rollout works on most requests and one API server still...A Gatekeeper policy rollout works on most requests and one API server still... focuses on admission-control and asks the reader to isolate the key signal in Kubernetes. Control-plane partial restarts can leave admission and conversion behavi...KubernetesAdvanced15 minProK8S-1367A kube-proxy migration to IPVS mostly works and one node blackholes service trafficA cluster networking optimization is rolled out and later only one node intermittently drops service traffic even though the migration reported success.KubernetesAdvanced15 minProK8S-1377A kube-proxy migration to IPVS mostly works and one node blackholes service trafficA network optimization rollout completes and later one node alone intermittently drops service traffic.KubernetesAdvanced15 minProK8S-1347A kube-proxy migration to IPVS passes and one service blackholes trafficA kube-proxy migration to IPVS passes and one service blackholes traffic focuses on network-segmentation and asks the reader to isolate the key signal in Kubernetes. Mode migrations can fail asymmetrically even when the daemonset rollout...KubernetesAdvanced15 minProK8S-1399A kube-proxy replacement migration works on worker nodes and one control...A kube-proxy replacement migration works on worker nodes and one control... focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Pod path success does not prove host-network exception beha...KubernetesAdvanced15 minProK8S-1357A kubeadm worker join passes and pods cannot pull imagesA kubeadm worker join passes and pods cannot pull images focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Successful node join does not prove the node runtime inherited the same registry assumptio...KubernetesAdvanced15 minProK8S-1345A kubelet certificate rotation succeeds and nodes still flip NotReadyA certificate maintenance window completes and afterward node readiness or metrics become inconsistent despite successful kubelet rotation.KubernetesAdvanced15 minProK8S-1353A kubelet eviction storm begins (eviction-threshold-hit-on-imagefs-not-dashboard-rootfs)A kubelet eviction storm begins (eviction-threshold-hit-on-imagefs-not... focuses on incident-response and asks the reader to isolate the key signal in Kubernetes. Storage health dashboards can be misleading if they do not align with kub...KubernetesAdvanced15 minProK8S-1325A projected service account token looks valid and a custom API aggregation...A projected service account token looks valid and a custom API aggregation... focuses on Identity And Access and asks the reader to isolate the key signal in Kubernetes. In Kubernetes auth, a valid token can still be unusable if its aud...KubernetesAdvanced15 minProK8S-1359A PVC snapshot restore looks complete and the app still corrupts writesA PVC snapshot restore looks complete and the app still corrupts writes focuses on incident-response and asks the reader to isolate the key signal in Kubernetes. Friendly snapshot names can mask wrong lineage in multi-tenant or repe...KubernetesAdvanced15 minProK8S-1384A Rancher-managed cluster upgrade completes and one node pool never rejoinsA Rancher-managed cluster upgrade completes and one node pool never rejoins focuses on Identity And Access and asks the reader to isolate the key signal in Kubernetes. Managed cluster reconnect failures often live in agent bootstra...KubernetesAdvanced15 minPro