Certification955 problems· 24 reviewed

CKA

955 incident response problems that help with CKA prep.

All problems (955)

K8S-1216Init container looks correct but still reads an empty secret fileA pod boots through an init step that reads a mounted secret and fails only on cold starts.KubernetesAdvanced18 minProK8S-116kube-proxy in IPVS mode retains a stale destination after a rapid rollout and some clients keep hitting terminated podsThe Service endpoints update correctly, yet a subset of traffic still lands on dead backends because the node-level forwarding table lagged behind the event stream.KubernetesAdvanced18 minProK8S-1224PDB allows disruption on paper but drain still stallsA node drain or maintenance window keeps failing even though the operator can count enough replicas in the cluster.KubernetesAdvanced18 minProK8S-1220Pod restarts stop after raising memory, but rollout still failsA team fixes the direct crash cause and still cannot complete rollout because the health timing contract was never updated.KubernetesAdvanced18 minProK8S-186Pod Security blocks the init-container workaround that the volume ownership model relied on during a failover rehearsalThe app and volume are healthy in isolation while the compatibility shim is no longer allowed to run. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced18 minProK8S-133The apiserver encryption provider order changes, and one controller can no longer read older Secrets until the keyring catches upCluster encryption is enabled, yet partial key visibility creates asymmetric control-plane failures.KubernetesAdvanced18 minProK8S-102Topology spread constraints look valid but matchLabelKeys exclude the new rollout revision and every new pod lands on one zoneThe policy appears healthy on paper, yet only the new ReplicaSet misses the balancing rule because its revision label is not part of the spread match set.KubernetesAdvanced18 minProK8S-1236A DaemonSet rollout looks healthy but one node class still failsA DaemonSet expands to a new node class and only that class starts crash-looping despite the same manifest working elsewhere.KubernetesAdvanced19 minProK8S-162A helper controller recreates pods faster than a node drain can evict them during a failover rehearsalMaintenance stalls because one control loop keeps undoing the intended disruption. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-276A kubelet eviction threshold watches imagefs while emptyDir growth is exhausting node ephemeral storage elsewhere during a failover rehearsalNode pressure is real while the selected signal points at the wrong filesystem. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-282A mutating webhook injects runtime labels that a later affinity rule needs, but scheduling decisions happen before those labels exist during a failover rehearsalThe final pod metadata looks correct while initial placement decisions were made without it. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-246A PodDisruptionBudget counts enough healthy pods while anti-affinity still prevents any replacement from landing during a failover rehearsalAvailability math looks safe until placement rules make recovery impossible. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-330A restore of etcd data succeeds while one API server still points at a stale peer list and keeps reintroducing dead endpoints into quorum logic during a failover rehearsalThe backing store is healthy, yet control-plane member configuration still trusts the old cluster map. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-152A restored etcd member rejoins with the old cluster ID still in its data dir, and control plane writes begin failing under quorum churnThe node is back online, but stale identity metadata makes the control plane unstable.KubernetesAdvanced19 minProK8S-222A StatefulSet ordinal comes back with a recycled PVC name while one sidecar still pins the old pod identity during a failover rehearsalStorage is healthy, yet one companion process keeps talking to the stateful member as if it were the previous ordinal instance. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-306A StatefulSet restore brings back volume data while one ordinal-specific peer identity still points at the pre-failover cluster member during a failover rehearsalThe data is back, but the peer relationship model remains attached to the old identity. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-115Audit log backend stalls and admission latency spikesRequests time out across the cluster, but the root cause is not the applications; it is the control plane waiting on an overloaded audit destination.KubernetesAdvanced19 minProK8S-1214Drain fails even with enough replicasAn operator drains a node and eviction keeps failing although replica count seems sufficient.KubernetesAdvanced19 minProK8S-1240Node drain keeps hangingA routine node maintenance window stalls when storage-backed workloads and their plugin surfaces cannot be drained in the expected order.KubernetesAdvanced19 minProK8S-1234PodDisruptionBudget blocks node maintenanceA cluster maintenance window starts and node drain cannot proceed because one replica has not been counted as available for a long time.KubernetesAdvanced19 minPro