Symptom384 problems· 9 reviewed

Rollout Stuck

384 incident problems that show up as “Rollout Stuck”.

All problems (384)

K8S-029CSI driver recovers volume attach only after controller restartPersistent volume attachment gets stuck until the controller is restarted, pointing to an unhealthy reconciliation loop.KubernetesAdvanced29 minProCICD-060Terraform PR plan is green but main apply uses a different variable setReviewers trust the plan output, but production apply behaves differently because the PR stage reads one workspace and tfvars path while the main branch apply uses another.CI/CDAdvanced29 minProNETWORK-130A NetFlow exporter uses a source interface in the management VRF, and the collector never receives records from production pathsTraffic forwarding is fine, but telemetry is dark because the export packets live in a different routing domain than the collector path.NetworkIntermediate15 minProCICD-003Adding an approval stage to an unverified production deployment pipelineA template-style problem for designing an approval flow and staging verification with per-environment deployment boundaries.CI/CDBeginner15 minProLINUX-393An application writes Unicode filenames correctly while a backup or sync service still runs under the old locale and drops part of the path set during a staged decommissionThe system stores the files and one scheduled path can no longer see them all. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.LinuxIntermediate15 minProLINUX-119An fstab automount hides the real mount failure until the first customer request hits the pathBoot looks clean, yet the outage begins later because the failing dependency is deferred behind on-demand mounting.LinuxIntermediate15 minProLINUX-102Journal rate limiting suppresses the real crash pattern and the team chases the wrong processOperators only see a handful of logs because the journal quota is dropping the repetitive fault signal during the outage window.LinuxIntermediate15 minProLINUX-092Kernel parameter is tuned in sysctl.d but an earlier file winsThe desired setting exists on disk, yet the host still uses another value because an earlier or later file in the sysctl load order overrides it.LinuxAdvanced15 minProCICD-120Release note automation rewrites the semantic version tag into a prerelease suffix and downstream promotion logic refuses the artifactThe published artifact exists, but the promotion controller rejects it because the computed version string no longer matches the expected stable pattern.CI/CDAdvanced15 minProLINUX-330A login shell exports one locale while the noninteractive service unit inherits another and decimal parsing diverges across the same host during a failover rehearsalNothing is wrong globally; the environment contract is simply different by execution mode. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.LinuxIntermediate16 minProCICD-210A release exception is approved while its audit evidence never reaches the archive during a failover rehearsalOperations can continue in the moment, yet the governance record for why it happened is incomplete. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.CI/CDIntermediate16 minProNETWORK-147A stack member replacement preserves the running config, but the trust device list for MACsec was stored in startup only and secure links never return after rebootOperations seem fine until the mandatory reboot reveals that a security dependency was not restored in both config states.NetworkIntermediate16 minProLINUX-121A systemd tmpfiles cleanup rule removes the idle service socket directory and the next connection starts failing long after bootThe host looks stable until the age-based cleanup task deletes a path the application expected to persist.LinuxAdvanced16 minProK8S-094CronJob concurrency policy skips every runThe schedule is correct, but no new work starts because the previous Job remains non-terminal and blocks the concurrency guard forever.KubernetesAdvanced16 minProNETWORK-128IPv6 ND inspection keeps a stale binding after renumbering, and hosts with new addresses cannot pass traffic on the access edgeSecurity stays enabled, but address lifecycle assumptions lag behind the renumbered endpoint state.NetworkIntermediate16 minProK8S-119Mutating webhook injects a sidecar into completed Jobs and the cleanup controller loops on already-finished podsThe platform team adds injection globally, but batch workloads begin failing cleanup because completed pods no longer match the assumptions of the controller.KubernetesAdvanced16 minProK8S-387A controller update passes admission while the stored object still carries an old finalizer contract that the new controller no longer clears during a staged decommissionCreate and update succeed, but lifecycle completion is stuck behind a legacy finalizer path. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.KubernetesAdvanced17 minProK8S-143A CSI snapshot restore succeeds, but the application pod still mounts the old PVC name through a leftover volumeClaimTemplate referenceData recovery completed, yet workload recovery still points at the pre-incident storage object.KubernetesAdvanced17 minProK8S-154A custom mutating webhook adds a sidecar to every pod, but the Jobs that already specify restartPolicy OnFailure now exceed the expected init sequenceThe platform-wide injection works for services, yet batch semantics drift because the startup contract changed.KubernetesAdvanced17 minProLINUX-143A dracut regeneration succeeds, but the host boots with the stale UEFI entry that still points at the previous initramfs pathThe recovery artifact exists, yet firmware is not loading it.LinuxAdvanced17 minPro