Symptom384 problems· 9 reviewed

Rollout Stuck

384 incident problems that show up as “Rollout Stuck”.

All problems (384)

CICD-155A blue-green switch updates the public ALB target group, but internal service discovery still resolves to the blue stack and background jobs keep writing thereThe front door moved cleanly, yet internal callers remain on the previous environment because they use a different discovery source.CI/CDAdvanced18 minProCICD-094Argo CD sync succeeds but a namespace label policy silently strips the network exemption labelThe app is deployed cleanly, yet the workload breaks because a cluster policy rewrites the namespace labels the app depends on for networking.CI/CDAdvanced18 minProLINUX-100Filesystem is remounted read-only after transient storage errors but the app hides it as generic 500sThe application logs are noisy, yet the real cause is the kernel remounting the filesystem read-only after I/O faults the app never surfaces clearly.LinuxAdvanced18 minProNETWORK-086IP SLA tracks the wrong probe target and withdraws the primary path during a healthy application incidentFailover triggers exactly as configured, but the chosen probe target does not represent the real application health and causes an unnecessary routing shift.NetworkAdvanced18 minProNETWORK-089OSPF adjacency forms but area type mismatch suppresses the LSAs needed by the branch routeNeighbor status looks correct, yet the expected route never appears because one side treats the area differently and filters critical LSA propagation.NetworkAdvanced18 minProCICD-204A blue-green cutover moves public ingress while internal service discovery stays on the previous stack during a failover rehearsalPublic traffic shifts cleanly but batch or backend callers continue writing into the old environment. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProCICD-079Mutable image tag makes an ECS rollback pull a newer build than the failed releaseThe rollback logic points at the previous task definition, yet the service still launches the wrong container because both revisions refer to the same mutable image tag.CI/CDAdvanced19 minProNETWORK-093OSPF adjacency remains full but the summary route hides the more specific path needed for policy routingThe protocol is healthy, yet the application path fails because route summarization obscures the specific destination the policy logic depends on.NetworkAdvanced19 minProCICD-091Release manifest points at a digest that exists only in the staging registryThe promotion metadata looks valid, but production cannot pull the image because the referenced digest was never replicated to the target registry.CI/CDAdvanced19 minProCICD-078Canary analysis passes against cached CDN responses instead of the origin traffic it was meant to judgeThe rollout looks healthy by metric, but the analysis is reading cache-hit behavior and not the new origin path that users will actually exercise after promotion.CI/CDAdvanced20 minProCICD-089Multi-account deployment role trusts the pipeline account but not the delegated tooling role session name patternCross-account deploys worked before, but now fail because the trust policy still allows the source account while denying the actual delegated session identity format.CI/CDAdvanced20 minProK8S-090Node reboot storm leaves CSI node plugin healthy but volume mounts failThe plugin pods appear up after recovery, but mounts still fail because kubelet is looking for a registration endpoint that the updated plugin no longer exposes in the same path.KubernetesAdvanced20 minProCICD-069OIDC deploy role works on push events but fails on workflow_call reuseThe repository already deploys successfully on direct pushes, but the reusable workflow path now fails because the token subject pattern no longer matches the calling context.CI/CDAdvanced20 minProLINUX-077Rsync restore with numeric-ids shifts file ownershipThe backup looks intact, but the restored application loses access because numeric UID preservation no longer matches the target system account layout.LinuxAdvanced20 minProLINUX-068tmpfs-backed runtime path fillsThe service outage looks like a simple restart loop, but the deeper issue is runtime storage pressure from large repeated coredumps on a limited tmpfs path.LinuxAdvanced20 minProK8S-086Cluster Autoscaler ignores pending podsPods stay pending and autoscaling never reacts because the requested local storage profile cannot fit any node shape in the expansion group.KubernetesAdvanced21 minProCICD-076Cross-region artifact replication lags and the disaster-recovery deploy uses a partial releaseThe promotion completed in the primary region, but the DR environment pulls an incomplete artifact set because replication finished for the manifest before every dependent object arrived.CI/CDAdvanced21 minProK8S-097CSI snapshot restore completes but the filesystem UUID collision confuses the bootstrap scriptStorage comes back online, yet the app still fails because the restored filesystem identity collides with a value the startup logic treats as unique.KubernetesAdvanced21 minProCICD-065ECR lifecycle cleanup deletes one architecture image and arm nodes start failing pullsThe repository still contains the expected tag, but multi-architecture pulls break on one platform because the manifest list points to a child image that was already expired.CI/CDAdvanced21 minProSECURITY-080Endpoint isolation policy blocks the EDR cloud callback and the host never recovers from containmentContainment starts correctly, but the host stays permanently isolated because the policy also cut off the control channel required to release it safely.SecurityAdvanced21 minPro