Certification955 problems· 24 reviewed

CKA

955 incident response problems that help with CKA prep.

All problems (955)

K8S-1317An HPA looks idle (metrics-adapter-cached-old-label-discovery)An HPA looks idle (the key signal) focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Visible metrics do not prove the autoscaler is querying them through the same cached path.KubernetesAdvanced15 minProK8S-1323An HPA sees no metrics after a label cleanupA platform team standardizes labels and later an HPA stops seeing metrics even though the backing monitoring query still returns data.KubernetesAdvanced15 minProK8S-1339An OpenSearch cluster behind Kubernetes ingress looks healthy and still drops...An OpenSearch cluster behind Kubernetes ingress looks healthy and still... focuses on Identity And Access and asks the reader to isolate the key signal in Kubernetes. Session loops can come from two valid cookies that disagree...KubernetesAdvanced15 minProK8S-1355An OpenSearch cluster on Kubernetes stays yellowA zone-label migration is rolled out and later stateful search replicas remain Pending even though enough nodes exist in the cluster.KubernetesAdvanced15 minProK8S-1395An OpenSearch operator rollout looks healthy and shard allocation still...An OpenSearch operator rollout looks healthy and shard allocation still... focuses on incident-response and asks the reader to isolate the key signal in Kubernetes. Shard churn can come from helper scripts reading pod annotations long afte...KubernetesAdvanced15 minProK8S-1343An OpenSearch operator rollout restarts pods foreverA chart rename seems harmless and later one custom resource keeps failing reconcile with admission or validation errors.KubernetesAdvanced15 minProK8S-1375An OpenSearch StatefulSet keeps a replica PendingA zone-label migration completes and later one search replica remains Pending despite enough nodes across zones.KubernetesAdvanced15 minProK8S-1365An OpenSearch StatefulSet keeps one replica PendingA label migration finishes and later one search replica remains unschedulable despite apparent cluster capacity.KubernetesAdvanced15 minProK8S-1341A Cilium cluster upgrade passes smoke tests and one namespace loses DNS egressA Cilium cluster upgrade passes smoke tests and one namespace loses DNS egress focuses on dns-resolution and asks the reader to isolate the key signal in Kubernetes. Name-based egress policy can fail when stale DNS state survives longer t...KubernetesAdvanced16 minProK8S-1331A Cilium policy rollout looks correct and one namespace still loses trafficA label cleanup is followed by namespace-specific traffic loss that appears only on a subset of nodes.KubernetesAdvanced16 minProK8S-1351A Cilium upgrade keeps policy green and one namespace loses east-west trafficA Cilium upgrade keeps policy green and one namespace loses east-west traffic focuses on cluster-maintenance and asks the reader to isolate the key signal in Kubernetes. Node-local dataplane state can outlive a healthy-looking Cilium r...KubernetesAdvanced16 minProK8S-1330A Loki deployment ingests logs correctly and recent queries look inconsistentA Loki deployment ingests logs correctly and recent queries look inconsistent focuses on incident-response and asks the reader to isolate the key signal in grafana. Recent-query inconsistency can be a convergence problem, not immediate...KubernetesAdvanced16 minProK8S-1312A service mesh sidecar is injected correctly and egress still failsA service mesh sidecar is injected correctly and egress still fails focuses on dns-resolution and asks the reader to isolate the key signal in Kubernetes. Sidecar policy debugging should include DNS path changes, not only network ACLs.KubernetesAdvanced16 minProK8S-1328A Traefik TCP route disappears only after a second team adds another serviceA cluster exposes multiple TCP services and one of them vanishes after another team deploys a new route on the same entryPoint.KubernetesAdvanced16 minProK8S-1319A Traefik TCP route works in staging and fails in productionA TCP ingress setup that works in staging misroutes in production where more routers and entryPoint reuse exist.KubernetesAdvanced16 minProK8S-1304A validating webhook only times out on node-local DNS nodesAdmission failures appear random after a service rename because only certain nodes run through a different DNS path.KubernetesAdvanced16 minProK8S-1321An admission webhook serves the right certificate and still times outAn admission webhook serves the right certificate and still times out focuses on network-segmentation and asks the reader to isolate the key signal in Kubernetes. Webhook reachability must be proven from the API server perspective, not only from r...KubernetesAdvanced16 minProK8S-1263An ALB target group stays unhealthyK8s incident scenario used for structured troubleshooting practice.KubernetesAdvanced16 minProK8S-1309An EKS add-on update looks clean but new pods cannot get IPsScaling or upgrading node groups suddenly reduces pod density while the network add-on still reports healthy.KubernetesAdvanced16 minProK8S-1314An ingress migration preserves the hostname and still failsAn ingress migration preserves the hostname and still fails focuses on Deployment Governance and asks the reader to isolate the key signal in Kubernetes. Controller migrations preserve intent less faithfully than YAML similarity suggests.KubernetesAdvanced16 minPro