Vendor339 problems· 9 reviewed

AWS

339 incident problems in AWS environments.

All problems (339)

CICD-148The artifact retention job keeps the container image but deletes the detached attestation blob, and production admission starts blocking the next rolloutBuild and publish succeeded, yet the downstream verifier requires a side artifact that retention policy treated as optional.CI/CDAdvanced17 minProCICD-351A blue-green cutover validates HTTP health while background lease holders still point at the retiring environment during a staged decommissionThe front door is healthy, but a hidden ownership path keeps mutating state from the wrong side. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.CI/CDAdvanced18 minProCICD-1289A canary deploy health check stays redA canary deployment begins failing only after a cluster or auth dependency upgrade, even though application code and rollout logic are unchanged.CI/CDAdvanced18 minProK8S-1289A cluster appears healthy until a node recycleA cluster runs for months and suddenly new nodes cannot create workloads because the admission webhook is unreachable only from fresh nodes.KubernetesAdvanced18 minProCICD-330A container build uses one CA bundle during image creation while the runtime base layer refreshes and trusts a different internal PKI root during a failover rehearsalThe built image and the later-executed image lineage no longer share the same trust store assumptions. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced18 minProCICD-1265A Docker build pushes successfully but later pulls failCI/CD incident scenario used for structured troubleshooting practice.CI/CDAdvanced18 minProCICD-1270A Docker build pushes successfully but later pulls failA registry migration leaves CI green while production deploys fail or pull an unexpected image because the runtime host is still logged into the former registry.CI/CDAdvanced18 minProK8S-1277A node group joins the cluster but DaemonSet pods stay brokenA custom AMI or launch template brings new nodes online and only platform add-ons behave strangely on those instances.KubernetesAdvanced18 minProK8S-1281A pod restarts after every successful deploymentA rollout works on some nodes and fails on others immediately after a private endpoint or dependency address changed.KubernetesAdvanced18 minProCICD-234A release candidate passes smoke tests in one region while traffic warmup happens against another region during a failover rehearsalValidation says the build is safe, but the workload that receives real traffic is not the one the tests exercised. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDIntermediate18 minProCICD-477A reusable workflow signs container images correctlyA reusable workflow signs container images correctly focuses on ci-cd-release-safety and asks the reader to isolate Permission Denied in GitHub. 실무에서는 ci-cd-release-safety 문제를 볼 때 실패 단계만 보지 말고 최근 변경, 이미지 태그, 시크릿 주입, 롤백 가능 여부를 먼저 함께 확인하는 편이 빠릅니다.CI/CDIntermediate18 minProK8S-1295A rollout stalls on only one zoneA cluster drain starts cleanly and one workload never reschedules despite enough total cluster capacity.KubernetesAdvanced18 minProK8S-1271A StatefulSet keeps rescheduling but pods stay PendingA StatefulSet keeps rescheduling but pods stay Pending focuses on storage-operations and asks the reader to isolate the key signal in AWS. A valid PVC and a healthy node pool can still deadlock if no single zone satisfies both scheduling and storage rules.KubernetesAdvanced18 minProK8S-1287A StatefulSet recovers after reboot but one member keeps attaching the wrong volumeA stateful workload is migrated or restored and one ordinal repeatedly boots with the wrong persistent data.KubernetesAdvanced18 minProCICD-147A Terraform destroy plan targets the right workspace, but the provider alias resolution shifted and the plan now points at the shared services accountReviewers looked at the expected stack, yet runtime provider wiring changed the real blast radius.CI/CDAdvanced18 minProCICD-128A Terraform plan file is generated with one provider plugin version and applied later with another, producing an unexpected drift errorThe change was reviewed correctly, but the saved plan no longer matches the provider behavior available at apply time.CI/CDAdvanced18 minProK8S-312An externalTrafficPolicy Local design is correctAn externalTrafficPolicy Local design is correct focuses on cluster-networking-and-service-discovery and asks the reader to isolate Timeouts and Latency in AWS. 실무에서는 timeouts-and-latency 증상만 보고 Pod 하나에 매달리지 말고 이벤트, 이전 로그, Service/Endpoint, 최근 배포 변경을 한 번에 묶어 보는 편이 오진을 줄입니다.KubernetesAdvanced18 minProCICD-312An IaC validation stage ignores tag-only drift while a compliance controller uses those tags to permit or deny later runtime changes during a failover rehearsalThe infrastructure shape matches expectations, yet its governance behavior does not. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced18 minProK8S-1259An ingress controller starts but health checks still failAn ingress controller starts but health checks still fail focuses on load-balancer-operations and asks the reader to isolate the key signal in AWS. Having the right tag keys is not the same as having subnets that are currently usable for the reque...KubernetesAdvanced18 minProCICD-240An integrity gate trusts an object-store ETag after transparent recompression changed the real payload identity during a failover rehearsalThe object is present and the hash-like signal no longer represents the original binary content. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced18 minPro