Symptom159 problems· 15 reviewed

Resource Exhaustion

159 incident problems that show up as “Resource Exhaustion”.

All problems (159)

LINUX-210Crash dumps remain enabled while the reserved crash kernel memory no longer matches the current kernel footprint during a failover rehearsalThe host thinks it is protected while the dump path quietly truncates or fails. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.LinuxAdvanced18 minProK8S-107Image filesystem eviction thresholds are configured, but node pressure is actually on rootfs and kubelet never evicts the bloated log directoryThe node shows pressure symptoms, yet eviction does not trigger because the exhausted partition is not the one covered by the active threshold.KubernetesAdvanced18 minProLINUX-111rsyslog disk-assisted forwarding queue fills the spool partition and local logging blocks behind remote delivery pressureRemote forwarding is supposed to be resilient, yet local observability degrades because the spillover queue shares space with critical system paths.LinuxAdvanced18 minProLINUX-104XFS project quota is enabled but a subdirectory inherits the wrong project ID and one tenant can overrun its limitQuota appears configured on the filesystem, yet nested paths escape the intended control because the directory tree was not fully assigned.LinuxAdvanced18 minProCICD-324A canary approval checks HTTP success while the job queue depth quietly rises behind the scenes and the release passes on the wrong health dimension during a failover rehearsalUser-facing requests look fine, but asynchronous backlog is already proving the release is unhealthy. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProK8S-276A kubelet eviction threshold watches imagefs while emptyDir growth is exhausting node ephemeral storage elsewhere during a failover rehearsalNode pressure is real while the selected signal points at the wrong filesystem. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProCICD-145A phased rollback restores traffic routing but leaves the message queue schema on the new version, and old workers poison retry trafficUser-facing paths look recovered, yet asynchronous paths still speak the newer contract and create delayed instability.CI/CDAdvanced19 minProCICD-180A rollback restores the app path but leaves an asynchronous contract on the newer version during a failover rehearsalUser-facing traffic recovers while queue consumers, workers, or webhooks still run with incompatible assumptions. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProLINUX-186A writeback cache hides a failing SSD until metadata finally flips read-only and makes recovery much harder during a failover rehearsalPerformance stays good while risk quietly accumulates behind the cache layer. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.LinuxAdvanced19 minProK8S-115Audit log backend stalls and admission latency spikesRequests time out across the cluster, but the root cause is not the applications; it is the control plane waiting on an overloaded audit destination.KubernetesAdvanced19 minProLINUX-095LVM thin pool reports plenty of virtual space but metadata exhaustion stops every writeThe logical volumes still look large enough, yet writes fail because the thin pool metadata area, not the data area, reached exhaustion first.LinuxAdvanced19 minProCICD-115Blue-green cutover updates ingress weights but the background cron target still writes into the old database schemaUser traffic looks healthy after the switch, yet scheduled jobs still point at the old stack and corrupt data alignment across environments.CI/CDAdvanced20 minProLINUX-028Container host shows low CPU while one cgroup throttles hardContainer host shows low CPU (Resource Exhaustion) is a hands-on troubleshooting drill. Overall node utilization looks safe, but one service is heavily throttled because its cgroup limit is too low. Linux Service Operations needs to be checked by narrowing scope, recent change...LinuxIntermediate20 minProK8S-083HPA scales on CPU but the true bottleneck is a single-threaded queue consumer partition assignmentAutoscaling appears to help, yet throughput never rises because the workload's shard or partition ownership keeps all useful work pinned to one replica.KubernetesAdvanced22 minProLINUX-073mdadm auto-assembly picks the wrong member order after an initramfs still references stale metadataThe disks are present, but recovery stays risky because the boot environment is assembling the array from outdated metadata assumptions.LinuxAdvanced23 minProK8S-061Topology spread and zone taints keep replicas pending after a node drainCapacity still exists in the cluster, but the workload cannot recover because spread rules and taints together eliminate every legal placement choice.KubernetesAdvanced24 minProK8S-037TopologySpreadConstraints leave replicas pending although one node has roomThe cluster has enough raw CPU and memory, but zone spread requirements and maxSkew prevent any legal placement for the remaining replicas.KubernetesAdvanced26 minProNETWORK-029Conntrack eviction hurts new connections before CPU looks busyConntrack eviction hurts new connections (Resource Exhaustion) is a hands-on troubleshooting drill. The host still has spare CPU, but new flows fail because connection tracking entries are evicted too aggressively. DNS and Routing needs to be checked by narrowing scope, recent...NetworkAdvanced29 minProLINUX-142A log shipping agent rotates its own state file to the root partition and eventually fills the boot disk instead of the intended data mountObservability remains intact, but the persistence path for agent state quietly lives on a much smaller filesystem.LinuxIntermediate15 minProCICD-135A rollback job restores the config map but not the worker scale, and the system keeps over-consuming the old backendConfiguration returns to normal, yet runtime pressure stays high because capacity settings were treated as outside the rollback scope.CI/CDIntermediate15 minPro