Cluster Maintenance
81 incident problems about Cluster Maintenance. Start with the reviewed ones.
먼저 읽을 가이드
추천 문제
All problems (81)
K8S-1214Drain fails even with enough replicasAn operator drains a node and eviction keeps failing although replica count seems sufficient.KubernetesAdvanced19 minProK8S-1400A cleanup job deletes old ReplicaSets and one rollback becomes impossibleA cleanup job deletes old ReplicaSets and one rollback becomes impossible focuses on Deployment Governance and asks the reader to isolate the key signal in Kubernetes. Operational rollback tooling often depends on annotations and histor...KubernetesIntermediate10 minProLINUX-1387A tmpfiles cleanup removes the app cache every bootA tmpfiles cleanup removes the app cache every boot focuses on cluster-maintenance and asks the reader to isolate the key signal in Linux. Path semantic changes can quietly bring an app under cleanup rules it was never designed to tolerate.LinuxIntermediate10 minProLINUX-1397A bridge interface works and one VLAN loses DHCPA host bridge cleanup modernizes VLAN handling and later one attached appliance stops receiving DHCP only after reboot.LinuxIntermediate11 minProLINUX-1383A dnf module stream migration succeeds on most hosts and one host keeps the old runtimeA fleet standardizes on a new runtime stream and one machine alone keeps reverting to the previous module set.LinuxIntermediate11 minProK8S-1390A PodDisruptionBudget is respected and a maintenance drain still never...A PodDisruptionBudget is respected and a maintenance drain still never... focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Successful eviction policy does not guarantee practical drain completion within your...KubernetesIntermediate11 minProLINUX-1391A systemd service starts manually and fails on bootA systemd service starts manually and fails on boot focuses on cluster-maintenance and asks the reader to isolate the key signal in Linux. Boot-time systemd conditions are one-shot gatekeepers unless something explicitly retriggers the...LinuxIntermediate11 minProLINUX-1357A package reinstall restores the binary and SELinux still blocks executionA recovery action reinstalls a package and later the service still fails even though the required binary now exists again.LinuxIntermediate12 minProLINUX-1361A systemd path unit stopped triggering after a deployment model switched from in-place writes to symlink swaps for the current release pointerThe watched path still changes logically, yet no event fires because the release tool now updates a symlink target instead of modifying the file directly.LinuxIntermediate12 minProLINUX-1341A systemd service works by hand and fails on bootA systemd service works by hand and fails on boot focuses on cluster-maintenance and asks the reader to isolate the key signal in Linux. Systemd may need files to exist before the process ever starts, not only by the time the command runs.LinuxIntermediate12 minProK8S-1398A Velero restore finishes and one application never comes upA Velero restore finishes and one application never comes up focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Restore plugins can quietly alter object content even when restore status looks successful.KubernetesIntermediate12 minProLINUX-1363An Ubuntu apt migration adds signed-by cleanly and one host still failsA repository trust migration is rolled out and later one host class keeps failing apt update despite the new signed-by config being present.LinuxIntermediate12 minProLINUX-1331A netplan change looks valid and the host still boots without networkA refreshed Ubuntu image is rolled out and only new instances lose network while the same config works on existing hosts.LinuxIntermediate13 minProK8S-1336A pod resolves services on old nodes and times out on fresh nodesA DNS service migration completes and later only pods on newly added nodes see timeouts or stale resolution results.KubernetesIntermediate13 minProK8S-1324A PodDisruptionBudget blocks every node drainA PodDisruptionBudget blocks every node drain focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Terminating is not the same as absent from the budget calculation.KubernetesIntermediate13 minProLINUX-1313A sudo-based maintenance timer works on one host and fails on anotherA sudo-based maintenance timer works on one host and fails on another focuses on Identity And Access and asks the reader to isolate the key signal in Linux. Policy fragments are code; ordering changes behavior even when the file set looks identical.LinuxIntermediate13 minProLINUX-1321A systemd PathUnit stops triggering after a release process moves from in-place writes to symlink swaps for current release selectionThe watched path appears to change, yet the unit never fires because the update method changed from file modification to symlink target replacement.LinuxIntermediate13 minProK8S-1294A custom metrics HPA reads values successfully but never scalesA workload is migrated between namespaces and autoscaling appears frozen even though the metric backend is healthy.KubernetesIntermediate14 minProK8S-1334A node drain hangs even though replicas are healthyA maintenance window begins and a single namespace blocks node drains with admission errors unrelated to application health.KubernetesIntermediate14 minProK8S-1329A Pod logs DNS timeouts only on fresh nodesA Pod logs DNS timeouts only on fresh nodes focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Node-specific DNS failures often point to bootstrap order rather than to cluster-wide resolver health.KubernetesIntermediate14 minPro