AWS
339 incident problems in AWS environments.
먼저 읽을 가이드
추천 문제
All problems (339)
K8S-1291A node group looks healthy but fresh pods hit image pull timeoutsA cluster scales to a new node group and only workloads landing there begin failing image pulls.KubernetesAdvanced17 minProK8S-1285A node group scales out but new pods still fail CNI allocationEKS nodes come up successfully and pods still hit IP allocation failures while another subnet in the VPC shows plenty of free space.KubernetesAdvanced17 minProK8S-1283A PodDisruptionBudget looks permissive enough but cluster upgrades still stallA managed cluster upgrade stalls on one workload even though operators believe the PodDisruptionBudget should allow at least one eviction.KubernetesAdvanced17 minProK8S-1293A PVC remains PendingA stateful set scales one more replica and the claim exists but the pod never reaches Running.KubernetesAdvanced17 minProCICD-393A release artifact is immutable while the runtime startup still downloads a policy pack from a bucket with newly narrowed cross-account access during a staged decommissionThe artifact did not change, but its boot dependency contract did. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.CI/CDAdvanced17 minProCICD-1281A reusable workflow can mint an OIDC token in one repository but fails in anotherA deployment pipeline is refactored into a shared workflow and later only one repository can no longer assume the cloud role.CI/CDAdvanced17 minProK8S-140After an upgrade, the kubelet image credential provider binary path changes, and pulls from the private registry fail on the new nodes onlyLegacy nodes keep working, but fresh workers cannot execute the helper that mints registry credentials.KubernetesAdvanced17 minProCICD-153An emergency patch bypasses the normal image scan stage, but admission still expects the scan metadata label and production refuses the deploymentThe fast path skips validation intentionally, yet the runtime guardrail still requires proof that only the normal path produces.CI/CDAdvanced17 minProK8S-1279An external-dns update looks successful but users still hit the wrong load balancerA cluster migration changes DNS and the automation pipeline reports success while real traffic continues to resolve the former endpoint.KubernetesAdvanced17 minProCICD-158An IaC drift detector ignores a tag-only difference, but the omitted tag controls an SCP exception and the next deploy fails in production onlyThe infrastructure looks equivalent in shape, yet one governance-relevant tag changed the allowed behavior.CI/CDAdvanced17 minProCICD-375An infrastructure drift policy ignores deleted optional resources while a recovery run still expects their outputs to exist for rollback wiring during a staged decommissionThe live state looks acceptable until rollback tries to consume outputs from deleted optional blocks. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.CI/CDAdvanced17 minProK8S-1299An ingress controller pod is healthy but its target group never updatesA controller chart is upgraded and later ingress state stops reconciling even though the pods remain healthy.KubernetesAdvanced17 minProK8S-1297An IRSA role works in one namespace but fails in anotherA chart upgrade lands and one namespace loses AWS API access while another using the same role keeps working.KubernetesAdvanced17 minProK8S-1273Cluster autoscaler sees pending pods but still refuses to scaleTainted workloads remain Pending even though more nodes should solve the problem and the autoscaler deployment appears healthy.KubernetesAdvanced17 minProK8S-1305Grafana Loki reads logs and still misses new alertsA log platform migrates to object storage and alerting becomes unreliable despite apparently valid credentials and buckets.KubernetesAdvanced17 minProCICD-1294Parallel OIDC jobs fail sporadicallyA repository changes many workflow files at once and suddenly several OIDC-enabled jobs begin failing randomly.CI/CDAdvanced17 minProK8S-117Projected service account token expires and the sidecar never reloads it so calls to the cloud API fail hours laterEverything works after startup, but long-lived pods lose access because one component reads the token once and never reopens the projected file.KubernetesAdvanced17 minProSECURITY-128S3 encryption enforcement is enabled, but a legacy multipart client omits the required KMS context and uploads start failing midstreamThe bucket policy is correct, yet one older client implementation cannot satisfy the newer encryption contract.SecurityAdvanced17 minProSECURITY-1204Temporary allowlist fixes a false positive but breaks the trust boundary by matching a broader source range than intendedA fast incident fix restores service, yet the allowlist rule now covers far more source space than the team originally intended to trust.SecurityIntermediate17 minProSECURITY-1199Temporary exception in the WAF lives foreverA public WAF tuning pattern recommends a quick allow rule. The exception fixes the incident, then quietly becomes part of the long-term exposure surface.SecurityIntermediate17 minPro