CI/CD Release Safety
105 incident problems about CI/CD Release Safety. Start with the reviewed ones.
먼저 읽을 가이드
추천 문제
All problems (105)
CICD-078Canary analysis passes against cached CDN responses instead of the origin traffic it was meant to judgeThe rollout looks healthy by metric, but the analysis is reading cache-hit behavior and not the new origin path that users will actually exercise after promotion.CI/CDAdvanced20 minProCICD-076Cross-region artifact replication lags and the disaster-recovery deploy uses a partial releaseThe promotion completed in the primary region, but the DR environment pulls an incomplete artifact set because replication finished for the manifest before every dependent object arrived.CI/CDAdvanced21 minProCICD-065ECR lifecycle cleanup deletes one architecture image and arm nodes start failing pullsThe repository still contains the expected tag, but multi-architecture pulls break on one platform because the manifest list points to a child image that was already expired.CI/CDAdvanced21 minProCICD-099Image signing succeeds in CI but the admission policy rejects the signature issuer after root rotationThe build publishes a signed image, yet cluster admission fails because the verifier still trusts the old signing root only.CI/CDAdvanced21 minProCICD-087Rollback Lambda succeeds but the Auto Scaling warm pool still serves the failed launch templateThe rollback path updates the group, yet instances continue to launch with the broken version because the warm capacity pool retained the earlier template state.CI/CDAdvanced21 minProCICD-085CloudFormation import succeeds but later drift repair wants to replace the manually retained resourceThe stack stabilizes after import, yet a future update becomes dangerous because the imported resource shape still differs from what the template assumes is replaceable.CI/CDAdvanced22 minProCICD-093Terraform drift fix replaces a subnet that still holds the canary target group routeThe plan appears corrective, but applying it would cut live traffic because one supposedly stale subnet still anchors an active canary path.CI/CDAdvanced22 minProCICD-073Terraform plan looks safe but apply recreates IAM roles after a for_each key renameNo obvious destructive change is noticed in review, yet apply replaces active roles because the stable key used by for_each changed during a refactor.CI/CDAdvanced22 minProCICD-071Argo CD prune deletes a shared secretThe sync itself succeeds, but a cleanup step removes a namespace-scoped secret that another workload still depends on because ownership boundaries were not encoded safely.CI/CDAdvanced23 minProCICD-080Feature flag migration step runs before the dependent schema change reaches every shardThe rollout script succeeds centrally, but one shard still serves the old schema and the new flag path begins calling a column that does not exist everywhere yet.CI/CDAdvanced23 minProCICD-083Schema migration succeeds on the writer but read replicas still serve incompatible shape to canary trafficThe migration log looks successful, yet the canary still fails because replica lag leaves part of the traffic reading the old schema path.CI/CDAdvanced23 minProCICD-063Terraform remote state lock survives a killed apply in a cross-account backendA failed apply no longer holds any active process, but every later run still stops on the lock because the backend cleanup path never completed across accounts.CI/CDAdvanced23 minProCICD-075Blue-green node group cutover drains the only log shipper before the replacement path is readyThe new nodes are healthy for the app, but operational visibility disappears because the drain order removed a cluster-wide DaemonSet before the replacement fleet was fully attached.CI/CDAdvanced24 minProCICD-090Canary bake time is shorter than the queue visibility window and failure signals arrive after promotionThe rollout passes its bake stage, but hidden worker failures appear only after the queue timeout window elapses, long after traffic has fully shifted.CI/CDAdvanced24 minProCICD-066CodeDeploy validation hook times outThe deployment itself is healthy, but the lifecycle validation keeps failing because the hook Lambda cannot reach the internal API it uses to prove readiness.CI/CDAdvanced24 minProCICD-357A reusable workflow signs container images correctly while downstream promotion retags an unsigned digest from a side repository during a staged decommissionSupply-chain controls protect the main path, but a side promotion lane bypasses the signed artifact. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.CI/CDIntermediate16 minProCICD-138A deployment circuit breaker watches the startup probe, but the one-off migration pod shares the same label and trips the release incorrectlyApplication pods are healthy, yet rollout halts because the failure budget includes a different workload type with a separate lifecycle.CI/CDAdvanced17 minProCICD-151A deployment freeze opens for one hour, but the delayed approval queue starts the rollout after the window closes and the policy engine revokes credentials mid-releaseEverything was approved at the right time, yet execution drifted outside the governance boundary.CI/CDAdvanced17 minProCICD-157A package signing key rotates in the build stage, but the downstream install test still trusts only the old fingerprint and rejects the freshly built repositoryThe package was published correctly, yet the validation environment pins a previous trust root.CI/CDAdvanced17 minProCICD-252A queued ChatOps rerun inherits stale authorization after the original release freeze window already closed during a failover rehearsalThe command looks approved from the chat log while execution happens under a no-longer-valid auth context. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced17 minPro