Resource Exhaustion
159 incident problems that show up as “Resource Exhaustion”.
Read first
When the ConfigMap changed but the Pod keeps using the old valuesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when values don't refresh after a ConfigMap update because of envFrom, subPath, volume projection, or rollout trigger differences.Kubernetes3 min readWhen CoreDNS is Running but only DNS lookups failAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the CoreDNS Pod looks healthy but failures occur in the kube-dns path, upstream, node-local-dns, or policy.Kubernetes3 min readWhat to check first when CrashLoopBackOff appearsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a Pod restarts repeatedly and the real cause is left in the previous logs and events.Kubernetes3 min read
Recommended problems
Reviewed problems first, then problems with detailed scenarios.
Self-hosted runner: No space left on device in the middle of buildsSelf-hosted runner: No space left on device in the middle of builds is a hands-on troubleshooting drill. Find what fills a long-lived runner's disk. GitHub Runner Hygiene needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는...ReviewedCI/CDBeginner3 minFreeK8S-030Ingress class mismatch sends traffic to the wrong controllerIngress class mismatch sends traffic to the wrong controller is a hands-on troubleshooting drill. Routing rules are valid, but the ingress object is reconciled by a different controller than the team expected. Kubernetes Ingress and Traffic needs to be checked by narrowing sco...ReviewedKubernetesIntermediate19 minFreeK8S-033PVC expansion succeeds in the API but filesystem size never changesStorageClass and claim expansion are enabled, yet the workload still hits a full disk because the filesystem inside the volume was never resized or remounted correctly.ReviewedKubernetesIntermediate22 minFreeK8S-036HPA stays idleMetrics look available and the pod is busy, but autoscaling does nothing because the target resource request needed for utilization math is missing.ReviewedKubernetesBeginner16 minFreePods get Evicted in bulkDozens of pods on two nodes show Evicted. The nodes report DiskPressure and one container is using 18Gi of ephemeral storage after debug logging was left on.ReviewedKubernetesIntermediate18 minFreeExit code 137 (OOMKilled): a container restarts under loadExit code 137 (OOMKilled): a container restarts is a hands-on troubleshooting drill. Learn what OOMKilled and exit code 137 mean. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는 CrashLoop...ReviewedKubernetesBeginner3 minFree
All problems (159)
No space left on device: the disk is full and files cannot be writtenNo space left on device: the disk is full and files cannot be written is a hands-on troubleshooting drill. Use df and du to find what filled the disk. Linux Storage and Filesystems needs to be checked by narrowing scope, recent change, and the current live signal before rollba...ReviewedLinuxBeginner3 minFreeCORS error: No 'Access-Control-Allow-Origin' header is presentCORS error: No 'Access-Control-Allow-Origin' header is present is a hands-on troubleshooting drill. Read the most common CORS error and decide which side has to change. WAF and AppSec Controls needs to be checked by narrowing scope, recent change, and the current live signal b...ReviewedSecurityBeginner3 minFreeExit code 137 (OOMKilled): a container restarts under loadExit code 137 (OOMKilled): a container restarts is a hands-on troubleshooting drill. Learn what OOMKilled and exit code 137 mean. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는 CrashLoop...ReviewedKubernetesBeginner3 minFreeToo many open files: new connections fail under loadToo many open files: new connections fail (Resource Exhaustion) is a hands-on troubleshooting drill. Check a process's open-file limit against its open descriptors. resource-management needs to be checked by narrowing scope, recent change, and the current live signal before ro...ReviewedLinuxBeginner3 minFreePending: new pods do not start after scaling outPending: new pods do not start (Resource Exhaustion) is a hands-on troubleshooting drill. Read a FailedScheduling event to see why a pod cannot be placed. Kubernetes Scheduling and Capacity needs to be checked by narrowing scope, recent change, and the current live signal befo...ReviewedKubernetesBeginner3 minFreeSelf-hosted runner: No space left on device in the middle of buildsSelf-hosted runner: No space left on device in the middle of builds is a hands-on troubleshooting drill. Find what fills a long-lived runner's disk. GitHub Runner Hygiene needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는...ReviewedCI/CDBeginner3 minFreeK8S-030Ingress class mismatch sends traffic to the wrong controllerIngress class mismatch sends traffic to the wrong controller is a hands-on troubleshooting drill. Routing rules are valid, but the ingress object is reconciled by a different controller than the team expected. Kubernetes Ingress and Traffic needs to be checked by narrowing sco...ReviewedKubernetesIntermediate19 minFreeLINUX-001An inode problem where files cannot be created even though disk space is freeA problem for identifying and organizing inode exhaustion, which is easy to miss when looking only at df output.ReviewedLinuxBeginner16 minFreeLINUX-033Deleted log file keeps disk fullOperators remove the giant log file, but the filesystem usage never drops because the service keeps writing to an already deleted file descriptor.ReviewedLinuxIntermediate18 minFreeK8S-033PVC expansion succeeds in the API but filesystem size never changesStorageClass and claim expansion are enabled, yet the workload still hits a full disk because the filesystem inside the volume was never resized or remounted correctly.ReviewedKubernetesIntermediate22 minFreeK8S-036HPA stays idleMetrics look available and the pod is busy, but autoscaling does nothing because the target resource request needed for utilization math is missing.ReviewedKubernetesBeginner16 minFreePods get Evicted in bulkDozens of pods on two nodes show Evicted. The nodes report DiskPressure and one container is using 18Gi of ephemeral storage after debug logging was left on.ReviewedKubernetesIntermediate18 minFreeThe OOM killer terminates the database during a nightly batchThe OOM killer terminates the database (Resource Exhaustion) is a hands-on troubleshooting drill. Read kernel OOM logs to find which process caused memory exhaustion and why the database was chosen. linux-performance-and-observability needs to be checked by narrowing scope, re...ReviewedLinuxIntermediate18 minFreeOOMKilled: a container keeps restarting after hitting its memory limitOOMKilled: a container keeps restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read OOMKilled and exit code 137, then fix memory limits and the workload that exceeds them. kubernetes-workload-reliability needs to be checked by narrowing scope, recent cha...ReviewedKubernetesBeginner15 minFreeToo many open files: a systemd service hits its file descriptor limitToo many open files: a systemd service hits its file descriptor limit is a hands-on troubleshooting drill. Find the real per-process limit of a systemd service instead of the shell ulimit. resource-management needs to be checked by narrowing scope, recent change, and the curre...ReviewedLinuxBeginner15 minFreeLINUX-008It looked like a memory leak, but it was really an orphan process that never terminatedAnalyzes a situation where a side worker outlives the main process and holds on to resources.LinuxIntermediate21 minFreeLINUX-002Narrowing a /var partition surge with du and lsofA hands-on disk analysis scenario that extends into a file-handle problem where space is not reclaimed after deletion.LinuxIntermediate24 minFreeK8S-019kubectl apply succeeds but the actual resource is created in a different namespacekubectl apply succeeds but the actual resource is created in a different... is a hands-on troubleshooting drill. A situation where the current context and the manifest namespace default differ, so it deploys to an unintended place. Kubernetes Config and Rollouts needs to be ch...KubernetesBeginner14 minFreeSECURITY-146A SIEM parser now splits IPv6 addresses and ports correctly, but one detection rule still assumes IPv4 colon counts and silently stops matchingThe data quality improved, yet one analytic depended on the previous broken representation.SecurityIntermediate16 minFreeSECURITY-192A parser becomes more correct while one detection silently depends on the old broken field shape during a failover rehearsalData quality improves and a rule built on yesterday's bug stops matching. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.SecurityIntermediate17 minFree