Incident response guides 31

Signals to check first, command examples, common misdiagnoses, and recovery order. After reading, practice on a problem about the same failure.

What to check first when CrashLoopBackOff appearsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a Pod restarts repeatedly and the real cause is left in the previous logs and events.Kubernetes3 min readThe registry and Secret to check first for ImagePullBackOffAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the image name looks correct but the pull fails because of tag, registry credential, or network policy.Kubernetes3 min readHow to isolate the root cause of ImagePullBackOffAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when registry auth, tag drift, node egress, and mirror policy all look like the same error.Kubernetes3 min readWhen the Pod is Running but only Service traffic failsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when Pod status looks healthy but the Service, Endpoint, NetworkPolicy, or Ingress path is broken.Kubernetes3 min readWhen CoreDNS is Running but only DNS lookups failAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the CoreDNS Pod looks healthy but failures occur in the kube-dns path, upstream, node-local-dns, or policy.Kubernetes3 min readHow to check Helm values merge and environment override driftAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when GitOps sync succeeds but the actual rendered manifest differs from the expected configuration.Kubernetes3 min readWhen the ConfigMap changed but the Pod keeps using the old valuesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when values don't refresh after a ConfigMap update because of envFrom, subPath, volume projection, or rollout trigger differences.Kubernetes3 min readHow to split Kubernetes failures along Pod, Service, and rollout boundariesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when you need to isolate whether a Kubernetes failure started in Pod status, Service discovery, controller, storage, or network.Kubernetes3 min readHow to split Linux failures by systemd, permission, and filesystem signalsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a service fails but you need to isolate whether the cause is a systemd unit, permission, inode, mount, or process.Linux3 min readWhen a command works manually but fails only under cronAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the same command succeeds in an interactive shell but fails only under cron or a non-interactive environment.Linux3 min readWhen permissions are fixed but only new directories are created with the wrong groupAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when existing file permissions are correct but group inheritance breaks for newly created files and directories.Linux3 min readWhen sudo works in the shell but fails under automationAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when only automation jobs fail because of TTY, sudoers, environment reset, or service user differences.Linux3 min readWhen GitHub Actions succeeds but only the rollout failsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the build logs look fine but the failure surfaces only on the target deploy cluster or runtime.CI/CD3 min readWhen the CI cache restores the wrong dependency graphAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when package.json, the lockfile, and artifacts no longer agree even after a cache hit.CI/CD3 min readWhen a self-hosted runner job stays stuck in queuedAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the runner appears registered but can't pick up jobs because of label, group, network, or token issues.CI/CD3 min readThe checking order when GitHub Actions succeeds but only the deploy failsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the workflow is green but the failure happens only in the rollout, promotion, or health check stage.CI/CD3 min readWhen firewalld looks open but connections keep getting blockedAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a port rule appears to exist but the connection fails because of zone, runtime/permanent drift, source binding, or an upstream firewall.Network3 min readHow to separate timeout and connection refused by network pathAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when DNS, route, firewall, proxy, and listener states all look like the same connection failure.Network3 min readWhen DNS changed but some clients still hit the old backendAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the DNS record is updated but resolver cache, HTTP/2 keepalive, or client pools hold on to the old target.Network3 min readHow to check x509 unknown authority and missing intermediate certificatesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when TLS verification fails because of chain, CA bundle, or trust store differences rather than the server certificate itself.Security3 min readHow to isolate an mTLS trust bundle mismatchAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the client certificate, server certificate, CA bundle, and sidecar reload timing differ and the handshake fails.Security3 min readHow to safely diagnose a WAF 403 false positiveAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when it looks like a permissions problem but a specific header, payload, or rule group is blocking legitimate requests.Security3 min readHow to split IAM, TLS, and WAF failures into a security operations flowAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when authentication, authorization, certificate, WAF, and audit log signals mix together and look like a single access failure.Security3 min readWhen GPU nodes exist but the workload keeps failing to scheduleA guide to separating resource requests, taint/toleration, node selector, and driver readiness first when GPU nodes are visible but Pods stay Pending.Kubernetes3 min readWhen the PersistentVolume reattaches but the rollout keeps stallingA guide to viewing attach state and workload readiness separately when the volume looks attached but Pod replacement and rollout keep stalling.Kubernetes3 min readWhen only install scripts or runners fail because of /tmp noexecA guide to separating the cases where noexec on /tmp makes only certain installers, self-hosted runners, or temp-file-based deploys fail.Linux3 min readWhen Preview passes but the secret is missing only under the production feature flagA guide to splitting secret scope and environment contract first when the preview environment is healthy but the secret path looks empty only the moment you turn on a production flag or release toggle.CI/CD3 min readWhy only one VLAN works and the rest fail on a native VLAN mismatchA search-oriented symptom guide covering native VLAN mismatch, trunk configuration drift, partial inter-VLAN failures, and Router-on-a-Stick confusion.Network3 min readA checklist for when OSPF adjacency won't formA search-oriented checklist guide covering OSPF neighbor down, stuck adjacency, route loss after failover, and partial path drift.Network3 min readHow to distinguish Router-on-a-Stick from SVI in practiceA search-oriented guide comparing Router-on-a-Stick and SVI by inter-VLAN routing, gateway location, and which point to trust in the logs during an incident.Network2 min readWhen login callbacks break due to proxy header differences between NGINX and AzureA search-oriented comparison guide covering reverse proxy header mismatches, secure callback failures, and proxy behavior differences between the NGINX and Azure paths.Security3 min read