Resource Management
26 incident problems about Resource Management. Start with the reviewed ones.
Read first
How to split Linux failures by systemd, permission, and filesystem signalsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a service fails but you need to isolate whether the cause is a systemd unit, permission, inode, mount, or process.Linux3 min readWhen sudo works in the shell but fails under automationAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when only automation jobs fail because of TTY, sudoers, environment reset, or service user differences.Linux3 min readWhen permissions are fixed but only new directories are created with the wrong groupAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when existing file permissions are correct but group inheritance breaks for newly created files and directories.Linux3 min read
Recommended problems
Reviewed problems first, then problems with detailed scenarios.
Too many open files: a systemd service hits its file descriptor limitToo many open files: a systemd service hits its file descriptor limit is a hands-on troubleshooting drill. Find the real per-process limit of a systemd service instead of the shell ulimit. resource-management needs to be checked by narrowing scope, recent change, and the curre...ReviewedLinuxBeginner15 minFreeToo many open files: new connections fail under loadToo many open files: new connections fail (Resource Exhaustion) is a hands-on troubleshooting drill. Check a process's open-file limit against its open descriptors. resource-management needs to be checked by narrowing scope, recent change, and the current live signal before ro...ReviewedLinuxBeginner3 minFree
All problems (26)
Too many open files: new connections fail under loadToo many open files: new connections fail (Resource Exhaustion) is a hands-on troubleshooting drill. Check a process's open-file limit against its open descriptors. resource-management needs to be checked by narrowing scope, recent change, and the current live signal before ro...ReviewedLinuxBeginner3 minFreeToo many open files: a systemd service hits its file descriptor limitToo many open files: a systemd service hits its file descriptor limit is a hands-on troubleshooting drill. Find the real per-process limit of a systemd service instead of the shell ulimit. resource-management needs to be checked by narrowing scope, recent change, and the curre...ReviewedLinuxBeginner15 minFreeK8S-1577A Prometheus remote-write queue drains and one shard still backpressuresOne remote-write shard continues to backpressure after tuning.KubernetesAdvanced10 minProK8S-1493A Kubernetes API priority and fairness change lands and kubelet heartbeats still starveNode heartbeats degrade only for fresh autoscaled nodes after API priority and fairness tuning.KubernetesAdvanced12 minProK8S-1503An API priority and fairness update lands and kubelet heartbeats still starveNode heartbeats degrade only for fresh autoscaled nodes after APF tuning.KubernetesAdvanced12 minProK8S-1602A custom metrics fix is correct and one HPA still reads zeroOne HPA still reads zero after custom metrics fixes.KubernetesIntermediate8 minProK8S-1592A metrics adapter fix lands and one HPA still stays flatOne HPA still stays flat after metrics adapter corrections.KubernetesIntermediate8 minProK8S-1582A metrics adapter query is corrected and one HPA still reads zeroOne HPA still reads zero after adapter query corrections.KubernetesIntermediate9 minProK8S-1528A Karpenter provisioner launches nodes and pods remain pendingPods stay pending on one scheduling path after a label schema migration.KubernetesIntermediate10 minProK8S-1518A Karpenter provisioner launches nodes and pods still stay pendingPods remain pending only on one scheduling path after a Karpenter label schema migration.KubernetesIntermediate10 minProK8S-1512A metrics adapter reports healthy and HPA still does not scaleAn HPA stalls only after scrape interval tuning in the metrics stack.KubernetesIntermediate10 minProK8S-1532A metrics adapter reports healthy and HPA still stays flatAutoscaling stops only after metrics scrape interval tuning.KubernetesIntermediate10 minProK8S-1522A metrics adapter reports healthy and the HPA still stays flatAutoscaling stops only after metrics scrape interval tuning.KubernetesIntermediate10 minProK8S-1562A Prometheus remote-write path backs up and the HPA still reads calm trafficHPA underreacts only when remote-write throttling begins.KubernetesIntermediate10 minProK8S-1552A Prometheus rule fires and the HPA still idlesAutoscaling stops only for low-volume workloads after recording rule tuning.KubernetesIntermediate10 minProK8S-1508An HPA sees demand and never scalesAutoscaling stops after adapter query-step tuning even though Prometheus shows rising demand.KubernetesIntermediate10 minProK8S-1489An HPA sees rising load and refuses to scaleAutoscaling stops only after a deployment rename and metrics adapter refresh that otherwise looks healthy.KubernetesIntermediate10 minProK8S-1498An HPA sees rising metrics and never scalesAutoscaling stops only after query step tuning on the metrics adapter and Prometheus backend.KubernetesIntermediate10 minProK8S-1529A Velero restore completes and fresh pods still blockNew pods are blocked only in restored namespaces after policy changes.KubernetesAdvanced11 minProK8S-1539A Velero restore completes and fresh pods still blockNew pods are blocked only in restored namespaces after policy changes.KubernetesAdvanced11 minPro