Incident response guides 31

Signals to check first, command examples, common misdiagnoses, and recovery order. After reading, practice on a problem about the same failure.

What to check first when CrashLoopBackOff appearsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a Pod restarts repeatedly and the real cause is left in the previous logs and events.Kubernetes3 min readThe registry and Secret to check first for ImagePullBackOffAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the image name looks correct but the pull fails because of tag, registry credential, or network policy.Kubernetes3 min readHow to isolate the root cause of ImagePullBackOffAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when registry auth, tag drift, node egress, and mirror policy all look like the same error.Kubernetes3 min readWhen the Pod is Running but only Service traffic failsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when Pod status looks healthy but the Service, Endpoint, NetworkPolicy, or Ingress path is broken.Kubernetes3 min readWhen CoreDNS is Running but only DNS lookups failAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the CoreDNS Pod looks healthy but failures occur in the kube-dns path, upstream, node-local-dns, or policy.Kubernetes3 min readHow to check Helm values merge and environment override driftAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when GitOps sync succeeds but the actual rendered manifest differs from the expected configuration.Kubernetes3 min readWhen the ConfigMap changed but the Pod keeps using the old valuesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when values don't refresh after a ConfigMap update because of envFrom, subPath, volume projection, or rollout trigger differences.Kubernetes3 min readHow to split Kubernetes failures along Pod, Service, and rollout boundariesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when you need to isolate whether a Kubernetes failure started in Pod status, Service discovery, controller, storage, or network.Kubernetes3 min readWhen GPU nodes exist but the workload keeps failing to scheduleA guide to separating resource requests, taint/toleration, node selector, and driver readiness first when GPU nodes are visible but Pods stay Pending.Kubernetes3 min readWhen the PersistentVolume reattaches but the rollout keeps stallingA guide to viewing attach state and workload readiness separately when the volume looks attached but Pod replacement and rollout keep stalling.Kubernetes3 min read