AWS
339 incident problems in AWS environments.
먼저 읽을 가이드
추천 문제
All problems (339)
CICD-091Release manifest points at a digest that exists only in the staging registryThe promotion metadata looks valid, but production cannot pull the image because the referenced digest was never replicated to the target registry.CI/CDAdvanced19 minProSECURITY-093Secrets rotation updates the database user but leaves a cached connection pool authenticating with the old passwordThe new credential is valid, yet outages continue because the application pool never discarded existing sessions that still reuse the previous password flow.SecurityAdvanced19 minProSECURITY-1185AWS security group looks correct but EKS API access still failsA team follows networking advice from public threads and proves the API endpoint is reachable. Access still fails because the cluster does not trust the caller identity.SecurityAdvanced20 minProCICD-078Canary analysis passes against cached CDN responses instead of the origin traffic it was meant to judgeThe rollout looks healthy by metric, but the analysis is reading cache-hit behavior and not the new origin path that users will actually exercise after promotion.CI/CDAdvanced20 minProSECURITY-072IAM permission boundary blocks emergency admin role assumption despite the attached allow policyThe incident role seems fully privileged, but assumption still fails because the boundary silently caps effective access below the attached policy intent.SecurityAdvanced20 minProSECURITY-083KMS grant allows encrypt but one rotated alias points the application to a key without decrypt permissionThe secret path still looks valid, yet runtime failures begin because the alias now resolves to a different key than the policy and grants were built for.SecurityAdvanced20 minProCICD-089Multi-account deployment role trusts the pipeline account but not the delegated tooling role session name patternCross-account deploys worked before, but now fail because the trust policy still allows the source account while denying the actual delegated session identity format.CI/CDAdvanced20 minProK8S-090Node reboot storm leaves CSI node plugin healthy but volume mounts failThe plugin pods appear up after recovery, but mounts still fail because kubelet is looking for a registration endpoint that the updated plugin no longer exposes in the same path.KubernetesAdvanced20 minProCICD-069OIDC deploy role works on push events but fails on workflow_call reuseThe repository already deploys successfully on direct pushes, but the reusable workflow path now fails because the token subject pattern no longer matches the calling context.CI/CDAdvanced20 minProSECURITY-1193Cloud audit agent looks healthy but stopped shipping logs after a service-account rotationThe agent process runs, yet the audit pipeline is effectively dark because the credential or permission path it uses no longer matches the rotated identity.SecurityAdvanced21 minProK8S-086Cluster Autoscaler ignores pending podsPods stay pending and autoscaling never reacts because the requested local storage profile cannot fit any node shape in the expansion group.KubernetesAdvanced21 minProCICD-076Cross-region artifact replication lags and the disaster-recovery deploy uses a partial releaseThe promotion completed in the primary region, but the DR environment pulls an incomplete artifact set because replication finished for the manifest before every dependent object arrived.CI/CDAdvanced21 minProK8S-097CSI snapshot restore completes but the filesystem UUID collision confuses the bootstrap scriptStorage comes back online, yet the app still fails because the restored filesystem identity collides with a value the startup logic treats as unique.KubernetesAdvanced21 minProCICD-065ECR lifecycle cleanup deletes one architecture image and arm nodes start failing pullsThe repository still contains the expected tag, but multi-architecture pulls break on one platform because the manifest list points to a child image that was already expired.CI/CDAdvanced21 minProSECURITY-1189IAM role rotation looks complete but long-lived pods keep using stale credentialsA cloud role was rotated and policy updated, yet running workloads continue to fail or over-permit because old credentials remain cached in process state.SecurityAdvanced21 minProCICD-087Rollback Lambda succeeds but the Auto Scaling warm pool still serves the failed launch templateThe rollback path updates the group, yet instances continue to launch with the broken version because the warm capacity pool retained the earlier template state.CI/CDAdvanced21 minProSECURITY-075An SCP allows the recovery service but blocks the dependent KMS decrypt call during restoreThe incident playbook launches correctly, but restore still fails because the organization policy forgot the downstream KMS permission the service actually needs.SecurityAdvanced22 minProCICD-085CloudFormation import succeeds but later drift repair wants to replace the manually retained resourceThe stack stabilizes after import, yet a future update becomes dangerous because the imported resource shape still differs from what the template assumes is replaceable.CI/CDAdvanced22 minProSECURITY-087CloudTrail organization trail exists but one delegated admin account writes to an unmonitored bucket in another regionAudit coverage seems complete, yet one privileged path is effectively invisible because the delegated admin is using a destination outside the monitored collection pattern.SecurityAdvanced22 minProSECURITY-031IAM role trust policy rejects GitHub OIDC tokenThe workflow reaches the cloud provider, but the trust policy denies the token because the expected audience or subject does not match the actual issuer claims.SecurityIntermediate22 minPro