AWS
339 incident problems in AWS environments.
먼저 읽을 가이드
추천 문제
All problems (339)
CICD-288A staged rollback restores old manifests while the backing secret reference already rotated to a new key family during a failover rehearsalThe old release comes back up, but its runtime secret contract no longer exists in the same form. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProK8S-1243A StatefulSet pod restarts forever after rescheduleA stateful workload reschedules after a node event and later fails repeatedly with permission errors on the data path.KubernetesAdvanced19 minProK8S-1253A StatefulSet volume reattaches after reschedule but write access still failsAfter a node failure, a stateful workload restarts on another node and immediately fails on file permissions despite the volume attaching cleanly.KubernetesAdvanced19 minProK8S-1248A StatefulSet volume reattaches successfully but startup still failsAfter a node event, a stateful workload comes back on another node and immediately fails with permission errors against an attached data volume.KubernetesAdvanced19 minProCICD-104Image promotion copies the manifest list but forgets the arm64 digest and only one node pool fails after releaseThe promoted tag looks valid in the registry, yet arm64 workers cannot pull because the multi-arch index lost a platform-specific child image during promotion.CI/CDAdvanced19 minProK8S-1229Image pulls fail only on one node groupA new node group joins and only workloads scheduled there begin showing image pull failures for a private registry.KubernetesAdvanced19 minProCICD-111Parallel Terraform plans race on the same remote state lock timeout and the second pipeline applies an outdated assumption setTwo changes target one workspace, one waits on the lock, and by the time it runs its plan is no longer trustworthy because the state already evolved.CI/CDAdvanced19 minProK8S-1256Pods in a private EKS subnet enter ImagePullBackOffA new private node group joins EKS and only workloads pulling from ECR begin failing despite valid image references and node IAM roles.KubernetesAdvanced19 minProK8S-1255A new node group becomes Ready but only pods there failA replacement or expansion node group joins the cluster and only workloads placed there begin failing on startup or networking operations.KubernetesAdvanced20 minProK8S-1250A node group becomes Ready but some workloads still failA fresh node group joins an EKS cluster and only pods landing there begin failing on startup or networking operations.KubernetesAdvanced20 minProK8S-1245A node group looks healthy but a subset of pods still cannot startA new node group joins an EKS cluster and only some workloads start failing after placement on the new instances.KubernetesAdvanced20 minProCICD-536A reusable workflow signs container images correctlyA reusable workflow signs container images correctly focuses on ci-cd-release-safety and asks the reader to isolate Permission Denied in GitHub. 실무에서는 ci-cd-release-safety 문제를 볼 때 실패 단계만 보지 말고 최근 변경, 이미지 태그, 시크릿 주입, 롤백 가능 여부를 먼저 함께 확인하는 편이 빠릅니다.CI/CDIntermediate20 minProK8S-1238An ingress route looks valid but one hostname still returns 404After adding or upgrading an ingress controller, one hostname starts returning the wrong response even though the Ingress object itself looks unchanged.KubernetesAdvanced20 minProK8S-084DaemonSet update stallsThe update strategy looks conservative, but rollout progress stops because the current unavailable state of the tainted group already exceeds the allowed update budget.KubernetesAdvanced20 minProCICD-096Rollback policy depends on ALB 5xx but the failure is only visible in gRPC status codesThe release is unhealthy, yet rollback never triggers because the monitored signal ignores application-layer failure semantics used by this service.CI/CDAdvanced20 minProK8S-1219ALB controller says permissions are correct but one action still failsAn ALB controller works enough to create some resources but repeatedly fails one reconciliation action.KubernetesAdvanced21 minProK8S-1221CoreDNS pods are Running but queries still failA cluster keeps CoreDNS in Running state after scaling, but name resolution fails only after pods land on a new node group.KubernetesAdvanced21 minProK8S-1208Managed node group reports join failure while nodes already appear Ready in the clusterThe console still shows a failure, but kubectl shows Ready nodes, sending the team into the wrong troubleshooting path.KubernetesAdvanced21 minProK8S-1227The ALB controller creates the load balancer but target registration failsAn Ingress object becomes active but downstream target health stays red after registration.KubernetesAdvanced21 minProCICD-039ECR lifecycle cleanup deletes digest still referenced by promotion manifestThe production promotion manifest points to an immutable digest, but registry cleanup removes the only copy before the delayed deployment window starts.CI/CDIntermediate22 minPro