Symptom384 problems· 9 reviewed

Rollout Stuck

384 incident problems that show up as “Rollout Stuck”.

All problems (384)

CICD-099Image signing succeeds in CI but the admission policy rejects the signature issuer after root rotationThe build publishes a signed image, yet cluster admission fails because the verifier still trusts the old signing root only.CI/CDAdvanced21 minProCICD-087Rollback Lambda succeeds but the Auto Scaling warm pool still serves the failed launch templateThe rollback path updates the group, yet instances continue to launch with the broken version because the warm capacity pool retained the earlier template state.CI/CDAdvanced21 minProCICD-085CloudFormation import succeeds but later drift repair wants to replace the manually retained resourceThe stack stabilizes after import, yet a future update becomes dangerous because the imported resource shape still differs from what the template assumes is replaceable.CI/CDAdvanced22 minProSECURITY-070EDR quarantine removes the log shipper binary and host visibility disappears without an alertThe endpoint agent did its job from one perspective, but security operations lose telemetry because the quarantined component was also the only path to central visibility.SecurityAdvanced22 minProNETWORK-090EVPN MAC mobility is detected but the dampening timer keeps traffic pinned to the old leafThe control plane sees the move, yet data still breaks because mobility dampening delays acceptance of the new location longer than the workload can tolerate.NetworkAdvanced22 minProK8S-068FailurePolicy Ignore lets pods start without the required security sidecarThe cluster stays available during webhook trouble, but production traffic later fails because workloads launched without the sidecar contract the platform assumes.KubernetesAdvanced22 minProK8S-028Node drain hangs on long-lived connection podsMaintenance starts correctly, but eviction never finishes because connection draining and termination hooks take too long.KubernetesIntermediate22 minProCICD-093Terraform drift fix replaces a subnet that still holds the canary target group routeThe plan appears corrective, but applying it would cut live traffic because one supposedly stale subnet still anchors an active canary path.CI/CDAdvanced22 minProCICD-073Terraform plan looks safe but apply recreates IAM roles after a for_each key renameNo obvious destructive change is noticed in review, yet apply replaces active roles because the stable key used by for_each changed during a refactor.CI/CDAdvanced22 minProK8S-076VolumeSnapshot restore binds to the wrong PVC lineage after a cloned recovery testThe snapshot data is valid, but a later restore attaches to the wrong expectation chain because snapshot content and clone naming were reused too casually during testing.KubernetesAdvanced22 minProLINUX-085XFS metadata corruption warning is stale but the on-call plans a destructive repair on the mounted volumeAn old alert resurfaces during an incident, and the real risk now is operator action because the filesystem is still mounted and serving production writes.LinuxAdvanced22 minProCICD-071Argo CD prune deletes a shared secretThe sync itself succeeds, but a cleanup step removes a namespace-scoped secret that another workload still depends on because ownership boundaries were not encoded safely.CI/CDAdvanced23 minProCICD-008Deployments passing without verificationA scenario that narrows down the root cause, centered on designing governance that enforces the verification step before deployment approval, in the situation of deployments passing without verification because a smoke-test conditional is wrong.CI/CDAdvanced23 minProCICD-080Feature flag migration step runs before the dependent schema change reaches every shardThe rollout script succeeds centrally, but one shard still serves the old schema and the new flag path begins calling a column that does not exist everywhere yet.CI/CDAdvanced23 minProCICD-083Schema migration succeeds on the writer but read replicas still serve incompatible shape to canary trafficThe migration log looks successful, yet the canary still fails because replica lag leaves part of the traffic reading the old schema path.CI/CDAdvanced23 minProK8S-062Stale VolumeAttachment object blocks PVC reattach after a node lossThe replacement node is ready, but the workload never mounts its volume because the storage control path still believes the old attachment is active.KubernetesAdvanced23 minProCICD-063Terraform remote state lock survives a killed apply in a cross-account backendA failed apply no longer holds any active process, but every later run still stops on the lock because the backend cleanup path never completed across accounts.CI/CDAdvanced23 minProCICD-075Blue-green node group cutover drains the only log shipper before the replacement path is readyThe new nodes are healthy for the app, but operational visibility disappears because the drain order removed a cluster-wide DaemonSet before the replacement fleet was fully attached.CI/CDAdvanced24 minProCICD-090Canary bake time is shorter than the queue visibility window and failure signals arrive after promotionThe rollout passes its bake stage, but hidden worker failures appear only after the queue timeout window elapses, long after traffic has fully shifted.CI/CDAdvanced24 minProCICD-066CodeDeploy validation hook times outThe deployment itself is healthy, but the lifecycle validation keeps failing because the hook Lambda cannot reach the internal API it uses to prove readiness.CI/CDAdvanced24 minPro