Symptom384 problems· 9 reviewed

Rollout Stuck

384 incident problems that show up as “Rollout Stuck”.

All problems (384)

NETWORK-114The multicast rendezvous point is reachable, but SSM clients failCore multicast signaling looks healthy, yet receivers never join the intended streams because the edge protocol expectation differs.NetworkAdvanced18 minProLINUX-174The right initramfs was built while firmware still boots an older entry pointing elsewhere during a failover rehearsalRecovery artifacts exist and the machine never actually loads them. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.LinuxAdvanced18 minProK8S-102Topology spread constraints look valid but matchLabelKeys exclude the new rollout revision and every new pod lands on one zoneThe policy appears healthy on paper, yet only the new ReplicaSet misses the balancing rule because its revision label is not part of the spread match set.KubernetesAdvanced18 minProCICD-324A canary approval checks HTTP success while the job queue depth quietly rises behind the scenes and the release passes on the wrong health dimension during a failover rehearsalUser-facing requests look fine, but asynchronous backlog is already proving the release is unhealthy. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProLINUX-294A dracut image includes the storage driver while an older UEFI entry still points at a kernel that lacks the matching userspace module set during a failover rehearsalEarly boot pieces are individually present and not aligned as one bootable pair. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.LinuxAdvanced19 minProCICD-216A final approval signs one manifest while the deploy job renders a different object set during a failover rehearsalReview and execution diverge because mutation still happens after signoff. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProK8S-162A helper controller recreates pods faster than a node drain can evict them during a failover rehearsalMaintenance stalls because one control loop keeps undoing the intended disruption. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProNETWORK-318A multicast RP pair is configured while one group range is filtered from MSDP state propagation and failover never restores those streams during a failover rehearsalRedundancy exists and one class of streams never inherits it. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.NetworkAdvanced19 minProK8S-282A mutating webhook injects runtime labels that a later affinity rule needs, but scheduling decisions happen before those labels exist during a failover rehearsalThe final pod metadata looks correct while initial placement decisions were made without it. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProCICD-145A phased rollback restores traffic routing but leaves the message queue schema on the new version, and old workers poison retry trafficUser-facing paths look recovered, yet asynchronous paths still speak the newer contract and create delayed instability.CI/CDAdvanced19 minProK8S-246A PodDisruptionBudget counts enough healthy pods while anti-affinity still prevents any replacement from landing during a failover rehearsalAvailability math looks safe until placement rules make recovery impossible. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-330A restore of etcd data succeeds while one API server still points at a stale peer list and keeps reintroducing dead endpoints into quorum logic during a failover rehearsalThe backing store is healthy, yet control-plane member configuration still trusts the old cluster map. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-152A restored etcd member rejoins with the old cluster ID still in its data dir, and control plane writes begin failing under quorum churnThe node is back online, but stale identity metadata makes the control plane unstable.KubernetesAdvanced19 minProCICD-258A rollback restores application code while feature migration flags remain enabled in the old environment during a failover rehearsalThe binary path moves backward and the runtime behavior still follows the newer feature contract. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProCICD-180A rollback restores the app path but leaves an asynchronous contract on the newer version during a failover rehearsalUser-facing traffic recovers while queue consumers, workers, or webhooks still run with incompatible assumptions. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProCICD-300A rollout job waits for one migration task while a second schema-affecting job starts from another workflow in parallel during a failover rehearsalThe database change path appears serialized, yet another automation lane mutates the same contract at the same time. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDAdvanced19 minProK8S-222A StatefulSet ordinal comes back with a recycled PVC name while one sidecar still pins the old pod identity during a failover rehearsalStorage is healthy, yet one companion process keeps talking to the stateful member as if it were the previous ordinal instance. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-306A StatefulSet restore brings back volume data while one ordinal-specific peer identity still points at the pre-failover cluster member during a failover rehearsalThe data is back, but the peer relationship model remains attached to the old identity. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProLINUX-691An application writes Unicode filenames correctlyAn application writes Unicode filenames correctly focuses on linux-performance-and-observability and asks the reader to isolate Rollout Stuck. 실무에서는 linux-performance-and-observability 문제를 볼 때 서비스 로그만 보지 말고 inode, 파일시스템 여유, 포트 점유, systemd 상태, 최근 패키지 변경까지 같이 확인해야 원인을 빨리 좁힐 수 있습니다.LinuxIntermediate20 minProCICD-115Blue-green cutover updates ingress weights but the background cron target still writes into the old database schemaUser traffic looks healthy after the switch, yet scheduled jobs still point at the old stack and corrupt data alignment across environments.CI/CDAdvanced20 minPro