Topic44 problems· 3 reviewed

Kubernetes Workload Reliability

44 incident problems about Kubernetes Workload Reliability. Start with the reviewed ones.

All problems (44)

K8S-107Image filesystem eviction thresholds are configured, but node pressure is actually on rootfs and kubelet never evicts the bloated log directoryThe node shows pressure symptoms, yet eviction does not trigger because the exhausted partition is not the one covered by the active threshold.KubernetesAdvanced18 minProK8S-102Topology spread constraints look valid but matchLabelKeys exclude the new rollout revision and every new pod lands on one zoneThe policy appears healthy on paper, yet only the new ReplicaSet misses the balancing rule because its revision label is not part of the spread match set.KubernetesAdvanced18 minProK8S-162A helper controller recreates pods faster than a node drain can evict them during a failover rehearsalMaintenance stalls because one control loop keeps undoing the intended disruption. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-276A kubelet eviction threshold watches imagefs while emptyDir growth is exhausting node ephemeral storage elsewhere during a failover rehearsalNode pressure is real while the selected signal points at the wrong filesystem. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-282A mutating webhook injects runtime labels that a later affinity rule needs, but scheduling decisions happen before those labels exist during a failover rehearsalThe final pod metadata looks correct while initial placement decisions were made without it. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-246A PodDisruptionBudget counts enough healthy pods while anti-affinity still prevents any replacement from landing during a failover rehearsalAvailability math looks safe until placement rules make recovery impossible. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-330A restore of etcd data succeeds while one API server still points at a stale peer list and keeps reintroducing dead endpoints into quorum logic during a failover rehearsalThe backing store is healthy, yet control-plane member configuration still trusts the old cluster map. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-152A restored etcd member rejoins with the old cluster ID still in its data dir, and control plane writes begin failing under quorum churnThe node is back online, but stale identity metadata makes the control plane unstable.KubernetesAdvanced19 minProK8S-222A StatefulSet ordinal comes back with a recycled PVC name while one sidecar still pins the old pod identity during a failover rehearsalStorage is healthy, yet one companion process keeps talking to the stateful member as if it were the previous ordinal instance. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-306A StatefulSet restore brings back volume data while one ordinal-specific peer identity still points at the pre-failover cluster member during a failover rehearsalThe data is back, but the peer relationship model remains attached to the old identity. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-115Audit log backend stalls and admission latency spikesRequests time out across the cluster, but the root cause is not the applications; it is the control plane waiting on an overloaded audit destination.KubernetesAdvanced19 minProK8S-1121A topology spread rule looks balanced (Rollout Stuck)A topology spread rule looks balanced (Rollout Stuck) focuses on kubernetes-workload-reliability and asks the reader to isolate Rollout Stuck. 실무에서는 Rollout Stuck 증상만 보고 Pod 하나에 매달리지 말고 이벤트, 이전 로그, Service/Endpoint, 최근 배포 변경을 한 번에 묶어 보는 편이 오진을 줄입니다.KubernetesAdvanced31 minProK8S-153A node selector matches the new GPU node pool label, but taints were added during bootstrap and the workload never lands thereThe scheduler finds the right nodes by label, yet another admission rule still excludes them.KubernetesIntermediate15 minProK8S-150A ConfigMap reload sidecar watches inode changes, but the projected volume updates by symlink swap and the app never reloadsConfiguration changes exist in the pod, yet the watcher logic does not observe the form of filesystem mutation Kubernetes actually uses.KubernetesIntermediate16 minProK8S-147A Deployment uses maxUnavailable zero, but node pressure evicts old pods anyway and the rollout briefly drops below the intended floorThe rollout strategy is conservative, yet external eviction pressure overrides the update assumptions.KubernetesIntermediate16 minProK8S-155An app depends on downward API labels, but the rollout controller changes the label key and the pod keeps reading an empty file without crashingThe deployment succeeds, yet the runtime configuration silently degrades because one metadata contract changed.KubernetesIntermediate16 minProK8S-270An inotify-based sidecar watches the wrong inode set after a projected ConfigMap update swaps symlinks during a failover rehearsalConfiguration changes are present while the watcher logic never sees the right file system events. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesIntermediate16 minProK8S-139HPA scale-down stabilization holds extra replicas during a rollout, and the Deployment budget never frees enough capacity for the next revisionNothing is obviously broken, but overlapping controller safety windows create a deadlock on available resources.KubernetesIntermediate16 minProK8S-318A ConfigMap hot reload works for the main container while a sidecar cached the previous schema and never reopens the mounted files during a failover rehearsalThe pod has new config, but not every process inside the pod has new understanding. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesIntermediate17 minProK8S-198A conservative rollout budget still loses availability when cluster pressure starts evicting old pods during a failover rehearsalDeployment math is correct, yet platform pressure changes the effective availability story. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesIntermediate17 minPro