Kubernetes Scheduling and Capacity
31 incident problems about Kubernetes Scheduling and Capacity. Start with the reviewed ones.
Read first
When the ConfigMap changed but the Pod keeps using the old valuesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when values don't refresh after a ConfigMap update because of envFrom, subPath, volume projection, or rollout trigger differences.Kubernetes3 min readWhen CoreDNS is Running but only DNS lookups failAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the CoreDNS Pod looks healthy but failures occur in the kube-dns path, upstream, node-local-dns, or policy.Kubernetes3 min readWhat to check first when CrashLoopBackOff appearsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a Pod restarts repeatedly and the real cause is left in the previous logs and events.Kubernetes3 min read
Recommended problems
Reviewed problems first, then problems with detailed scenarios.
K8S-008New workloads not schedulingA scenario that narrows down the root cause, centered on comparing the schedule conditions with the actual node labels, in the situation of new workloads not scheduling because node affinity is too strong.ReviewedKubernetesIntermediate21 minFreeK8S-030Ingress class mismatch sends traffic to the wrong controllerIngress class mismatch sends traffic to the wrong controller is a hands-on troubleshooting drill. Routing rules are valid, but the ingress object is reconciled by a different controller than the team expected. Kubernetes Ingress and Traffic needs to be checked by narrowing sco...ReviewedKubernetesIntermediate19 minFreeK8S-036HPA stays idleMetrics look available and the pod is busy, but autoscaling does nothing because the target resource request needed for utilization math is missing.ReviewedKubernetesBeginner16 minFreePods get Evicted in bulkDozens of pods on two nodes show Evicted. The nodes report DiskPressure and one container is using 18Gi of ephemeral storage after debug logging was left on.ReviewedKubernetesIntermediate18 minFreePending: new pods do not start after scaling outPending: new pods do not start (Resource Exhaustion) is a hands-on troubleshooting drill. Read a FailedScheduling event to see why a pod cannot be placed. Kubernetes Scheduling and Capacity needs to be checked by narrowing scope, recent change, and the current live signal befo...ReviewedKubernetesBeginner3 minFreeK8S-007A metrics pipeline problem where scaling does not rise even though an HPA is attachedTracks the metrics flow in a situation where resource usage is high but the autoscaler does not react.KubernetesAdvanced26 minPro
All problems (31)
Pending: new pods do not start after scaling outPending: new pods do not start (Resource Exhaustion) is a hands-on troubleshooting drill. Read a FailedScheduling event to see why a pod cannot be placed. Kubernetes Scheduling and Capacity needs to be checked by narrowing scope, recent change, and the current live signal befo...ReviewedKubernetesBeginner3 minFreeK8S-030Ingress class mismatch sends traffic to the wrong controllerIngress class mismatch sends traffic to the wrong controller is a hands-on troubleshooting drill. Routing rules are valid, but the ingress object is reconciled by a different controller than the team expected. Kubernetes Ingress and Traffic needs to be checked by narrowing sco...ReviewedKubernetesIntermediate19 minFreeK8S-008New workloads not schedulingA scenario that narrows down the root cause, centered on comparing the schedule conditions with the actual node labels, in the situation of new workloads not scheduling because node affinity is too strong.ReviewedKubernetesIntermediate21 minFreeK8S-036HPA stays idleMetrics look available and the pod is busy, but autoscaling does nothing because the target resource request needed for utilization math is missing.ReviewedKubernetesBeginner16 minFreePods get Evicted in bulkDozens of pods on two nodes show Evicted. The nodes report DiskPressure and one container is using 18Gi of ephemeral storage after debug logging was left on.ReviewedKubernetesIntermediate18 minFreeK8S-021DaemonSet rolls out everywhere except tainted nodesDaemonSet rolls out everywhere except tainted nodes is a hands-on troubleshooting drill. The manifest looks correct, but taints and tolerations keep the agent from reaching the nodes that matter most. Kubernetes Scheduling and Capacity needs to be checked by narrowing scope, r...KubernetesBeginner16 minFreeK8S-018GPU workloads stay Pending even though the cluster autoscaler is onGPU workloads stay Pending even though the cluster autoscaler is on is a hands-on troubleshooting drill. A situation where the autoscaler is working but the workload waits because of a special node-group condition. Kubernetes Scheduling and Capacity needs to be checked by narr...KubernetesAdvanced29 minProK8S-007A metrics pipeline problem where scaling does not rise even though an HPA is attachedTracks the metrics flow in a situation where resource usage is high but the autoscaler does not react.KubernetesAdvanced26 minProK8S-010A StatefulSet PVC stuck Pending, stalling the rolloutA problem where a stateful workload does not come up because StorageClass, access mode, and capacity conditions do not match.KubernetesAdvanced28 minProK8S-090Node reboot storm leaves CSI node plugin healthy but volume mounts failThe plugin pods appear up after recovery, but mounts still fail because kubelet is looking for a registration endpoint that the updated plugin no longer exposes in the same path.KubernetesAdvanced20 minProK8S-086Cluster Autoscaler ignores pending podsPods stay pending and autoscaling never reacts because the requested local storage profile cannot fit any node shape in the expansion group.KubernetesAdvanced21 minProK8S-097CSI snapshot restore completes but the filesystem UUID collision confuses the bootstrap scriptStorage comes back online, yet the app still fails because the restored filesystem identity collides with a value the startup logic treats as unique.KubernetesAdvanced21 minProK8S-028Node drain hangs on long-lived connection podsMaintenance starts correctly, but eviction never finishes because connection draining and termination hooks take too long.KubernetesIntermediate22 minProK8S-076VolumeSnapshot restore binds to the wrong PVC lineage after a cloned recovery testThe snapshot data is valid, but a later restore attaches to the wrong expectation chain because snapshot content and clone naming were reused too casually during testing.KubernetesAdvanced22 minProK8S-062Stale VolumeAttachment object blocks PVC reattach after a node lossThe replacement node is ready, but the workload never mounts its volume because the storage control path still believes the old attachment is active.KubernetesAdvanced23 minProK8S-069Cross-node gRPC calls failIntra-node traffic is clean, but larger cross-node requests hang because one node pool uses a smaller effective MTU than the other path expects.KubernetesAdvanced24 minProK8S-071PodDisruptionBudget and topology spread leave no legal recovery layout after an AZ outageThe workload had enough replicas before the outage, but after one zone disappears the remaining rules make every recovery option violate either spread or disruption guarantees.KubernetesAdvanced24 minProK8S-073Ephemeral-storage eviction startsThe nodes still have CPU and memory, but pods get evicted because local ephemeral storage fills when retry-heavy request logging lands in emptyDir volumes.KubernetesAdvanced19 minProK8S-084DaemonSet update stallsThe update strategy looks conservative, but rollout progress stops because the current unavailable state of the tainted group already exceeds the allowed update budget.KubernetesAdvanced20 minProK8S-095Pod anti-affinity keeps a restore blockedCapacity exists, but strict anti-affinity rules stop recovery because the surviving node already runs the same workload family.KubernetesAdvanced20 minPro