Health Probes
12 incident problems about Health Probes. Start with the reviewed ones.
Read first
When the ConfigMap changed but the Pod keeps using the old valuesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when values don't refresh after a ConfigMap update because of envFrom, subPath, volume projection, or rollout trigger differences.Kubernetes3 min readWhen CoreDNS is Running but only DNS lookups failAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the CoreDNS Pod looks healthy but failures occur in the kube-dns path, upstream, node-local-dns, or policy.Kubernetes3 min readWhat to check first when CrashLoopBackOff appearsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a Pod restarts repeatedly and the real cause is left in the previous logs and events.Kubernetes3 min read
Recommended problems
Reviewed problems first, then problems with detailed scenarios.
All problems (12)
K8S-1192Core application restarts endlesslyA public probe example was copied into production. It worked in steady state, but after a cold deploy the app now gets killed before it can finish bootstrap.ReviewedKubernetesAdvanced21 minFreeK8S-1220Pod restarts stop after raising memory, but rollout still failsA team fixes the direct crash cause and still cannot complete rollout because the health timing contract was never updated.KubernetesAdvanced18 minProK8S-1234PodDisruptionBudget blocks node maintenanceA cluster maintenance window starts and node drain cannot proceed because one replica has not been counted as available for a long time.KubernetesAdvanced19 minProK8S-1244A readiness probe path exists but still failsA service mesh or auth sidecar is introduced and previously healthy readiness probes begin failing despite the app route being up.KubernetesIntermediate17 minProK8S-1249A readiness route exists but still failsA service mesh, auth proxy, or sidecar is introduced and previously healthy readiness checks begin failing despite the app route working normally for real clients.KubernetesIntermediate17 minProK8S-1254A readiness route exists but still failsAfter introducing a service mesh or auth proxy, previously healthy readiness checks begin failing despite the application still serving the route.KubernetesIntermediate17 minProK8S-1207CrashLoopBackOff debugging misses the real stack traceA rollout hits CrashLoopBackOff and on-call engineers keep tailing logs from the current container only.KubernetesIntermediate17 minProK8S-1204kubectl logs shows the symptom but not the causeA restart loop looks mysterious because current logs contain only probe or wrapper noise. The evidence that matters is in the previous instance that operators never inspected.KubernetesIntermediate17 minProK8S-1218Readiness remains false foreverA deployment copied a probe path from an older chart revision where an init step created the endpoint first.KubernetesIntermediate17 minProK8S-1211Readiness never stabilizesA service keeps flapping during boot and the team only increases delay seconds.KubernetesIntermediate18 minProK8S-1225Startup probe keeps failingA workload with a proxy sidecar fails startup probes even though the application container itself is healthy.KubernetesIntermediate18 minProK8S-1184Service has healthy pods but zero usable endpointsA public answer suggests checking selectors first. Selectors are correct, but readiness is the real gate keeping endpoints empty.KubernetesIntermediate20 minPro