Topic81 problems· 1 reviewed

Cluster Maintenance

81 incident problems about Cluster Maintenance. Start with the reviewed ones.

All problems (81)

K8S-1185PodDisruptionBudget blocks node drainA public hardening guide recommended a conservative PDB. Later, replica counts dropped, but the old minAvailable still blocks every drain.KubernetesAdvanced23 minProK8S-1200Controller pod runs but never actsA controller upgrade follows public guidance. The pod restarts cleanly, but reconciliations stop because the leader-election path cannot update its lease object.KubernetesAdvanced24 minProK8S-1205Controller starts fine after upgrade but CRD drift makes new objects fail validation on only one clusterThe controller image upgrade succeeds, but one cluster still rejects new objects because its CRD version or schema was not updated along with the controller.KubernetesAdvanced24 minProK8S-1195Draining one node silently breaks a local-path workload that was never truly portableA maintenance drain appears routine, but one workload fails badly because it depended on local storage behavior that was never made portable across nodes.KubernetesAdvanced24 minProNETWORK-1379A HA sync succeeds and a package upgrade still breaks TLSAn HA pair survives a software upgrade and later failover lands on a node that cannot present the expected certificate chain.NetworkIntermediate12 minProNETWORK-1346A Netgate HA pair fails over cleanly and one package VPN stays downAn HA firewall upgrade is validated by failover tests and later one VPN tunnel fails only on the secondary node.NetworkIntermediate13 minProLINUX-1330A package reinstall restores the binary and SELinux still denies executionA recovery action reinstalls a package and afterward one service cannot execute a script or binary that used to work before the incident.LinuxIntermediate13 minProNETWORK-1340A pfSense HA pair passes XMLRPC sync and still divergesAn HA firewall pair is upgraded and later one node behaves differently despite successful XMLRPC config sync.NetworkIntermediate13 minProK8S-1382A cert-manager CA injector updates one namespace and misses anotherA cert-manager CA injector updates one namespace and misses another focuses on Identity And Access and asks the reader to isolate the key signal in Kubernetes. Injection drift can come from label schema changes that silently drop one object...KubernetesAdvanced14 minProK8S-1397A Cilium FQDN policy allows package mirrors and node bootstrap still failsA Cilium FQDN policy allows package mirrors and node bootstrap still fails focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. FQDN policy allows are only effective after the observation path has se...KubernetesAdvanced14 minProLINUX-1382A cloud-init host boots cleanly and one app misses its secretsA cloud-init host boots cleanly and one app misses its secrets focuses on cluster-maintenance and asks the reader to isolate the key signal in ubuntu. One-time bootstrap failures often come from stage ordering and restarted dependencies,...LinuxAdvanced14 minProLINUX-1350A local package mirror update succeeds and apt still rejects metadataA local package mirror update succeeds and apt still rejects metadata focuses on Identity And Access and asks the reader to isolate the key signal in Linux. Package trust failures after migration often mean one host still uses a d...LinuxAdvanced14 minProLINUX-1386A NetworkManager bridge migration works and one VLAN loses connectivityA host networking migration moves to keyfiles and later one VLAN becomes intermittently unreachable after reboot.LinuxAdvanced14 minProLINUX-1360A package mirror migration looks complete and one host still rejects metadataA package mirror migration looks complete and one host still rejects metadata focuses on Identity And Access and asks the reader to isolate the key signal in Linux. Repository trust failures after migration often mean a host is s...LinuxAdvanced14 minProK8S-1310Grafana Agent remote_write looks connected but drops a burst of cluster metricsMetrics ingestion resumes after pressure recovery, but dashboards never fill the missing historical gap from the outage window.KubernetesIntermediate14 minProK8S-1306Promtail tails files successfully until log rotation acceleratesA noisy cluster starts losing short-lived container logs even though Promtail itself reports no crash or auth error.KubernetesIntermediate14 minProK8S-1298A CronJob overlaps unexpectedlyA scheduled batch task starts producing duplicate side effects during node churn or maintenance windows.KubernetesIntermediate15 minProK8S-1316A mutating webhook seems available and pod creation still slows to a crawlA mutating webhook seems available and pod creation still slows to a crawl focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Fresh endpoints do not guarantee fresh dataplane state on every node.KubernetesIntermediate15 minProK8S-1311A StatefulSet rollout stalls only on replacement nodesA storage add-on update seems fine until stateful workloads land on new nodes and never attach volumes successfully.KubernetesIntermediate15 minProK8S-1224PDB allows disruption on paper but drain still stallsA node drain or maintenance window keeps failing even though the operator can count enough replicas in the cluster.KubernetesAdvanced18 minPro