Vendor339 problems· 9 reviewed

AWS

339 incident problems in AWS environments.

All problems (339)

K8S-087OIDC provider issuer URL rotates and every projected token verifier in the cluster rejects new tokensToken projection still works, but consumers fail because the issuer trust path and JWKS discovery URL changed underneath long-lived verifiers.KubernetesAdvanced22 minProCICD-093Terraform drift fix replaces a subnet that still holds the canary target group routeThe plan appears corrective, but applying it would cut live traffic because one supposedly stale subnet still anchors an active canary path.CI/CDAdvanced22 minProCICD-073Terraform plan looks safe but apply recreates IAM roles after a for_each key renameNo obvious destructive change is noticed in review, yet apply replaces active roles because the stable key used by for_each changed during a refactor.CI/CDAdvanced22 minProK8S-076VolumeSnapshot restore binds to the wrong PVC lineage after a cloned recovery testThe snapshot data is valid, but a later restore attaches to the wrong expectation chain because snapshot content and clone naming were reused too casually during testing.KubernetesAdvanced22 minProSECURITY-1191WAF allows the obvious route but still misses the attackA WAF rollout follows public guidance and appears healthy. Later, testing shows the same app is still reachable through an alternate hostname that bypasses the protected edge path.SecurityAdvanced22 minProSECURITY-1183AWS WAF blocks a safe admin pathA team enabled an AWS managed rule set following public guidance. An internal admin path now breaks because its request pattern trips a generic rule.SecurityAdvanced23 minProSECURITY-1198Cloud audit trail is complete in one account but gaps appearA multi-account logging design followed public cloud guidance. A permission change in one account now creates partial visibility loss that is easy to miss from the central console.SecurityAdvanced23 minProSECURITY-1195Cloud role assumption succeeds in one region and fails in anotherA cloud auth rollout followed community guidance and worked in one region. A second region fails because the trust condition still expects the previous audience or issuer pattern.SecurityAdvanced23 minProSECURITY-053CSP nonce is generated correctly but disappears after CDN template cachingThe application renders a fresh nonce, yet the browser still blocks the script because the cached edge fragment reuses a stale header-body combination.SecurityAdvanced23 minProSECURITY-063KMS policy lets backup jobs encrypt but restore jobs cannot decrypt in the recovery accountBackups complete successfully, yet every restore attempt fails because the disaster-recovery account was never granted the full decrypt path for the same key.SecurityAdvanced23 minProCICD-083Schema migration succeeds on the writer but read replicas still serve incompatible shape to canary trafficThe migration log looks successful, yet the canary still fails because replica lag leaves part of the traffic reading the old schema path.CI/CDAdvanced23 minProK8S-062Stale VolumeAttachment object blocks PVC reattach after a node lossThe replacement node is ready, but the workload never mounts its volume because the storage control path still believes the old attachment is active.KubernetesAdvanced23 minProCICD-063Terraform remote state lock survives a killed apply in a cross-account backendA failed apply no longer holds any active process, but every later run still stops on the lock because the backend cleanup path never completed across accounts.CI/CDAdvanced23 minProCICD-075Blue-green node group cutover drains the only log shipper before the replacement path is readyThe new nodes are healthy for the app, but operational visibility disappears because the drain order removed a cluster-wide DaemonSet before the replacement fleet was fully attached.CI/CDAdvanced24 minProCICD-090Canary bake time is shorter than the queue visibility window and failure signals arrive after promotionThe rollout passes its bake stage, but hidden worker failures appear only after the queue timeout window elapses, long after traffic has fully shifted.CI/CDAdvanced24 minProSECURITY-038CDN caches an authenticated error pageThe login path itself is correct, but one personalized failure response gets cached at the edge and leaks confusing content to later users.SecurityAdvanced24 minProCICD-066CodeDeploy validation hook times outThe deployment itself is healthy, but the lifecycle validation keeps failing because the hook Lambda cannot reach the internal API it uses to prove readiness.CI/CDAdvanced24 minProK8S-069Cross-node gRPC calls failIntra-node traffic is clean, but larger cross-node requests hang because one node pool uses a smaller effective MTU than the other path expects.KubernetesAdvanced24 minProSECURITY-069Organization-wide CloudTrail is enabled but one region never uses the expected KMS keyAudit logging exists everywhere, yet one region violates the encryption standard because replication and key policy assumptions drifted apart over time.SecurityAdvanced24 minProK8S-071PodDisruptionBudget and topology spread leave no legal recovery layout after an AZ outageThe workload had enough replicas before the outage, but after one zone disappears the remaining rules make every recovery option violate either spread or disruption guarantees.KubernetesAdvanced24 minPro