Topic204 problems· 5 reviewed

Kubernetes Config and Rollouts

204 incident problems about Kubernetes Config and Rollouts. Start with the reviewed ones.

All problems (204)

CICD-067Argo Rollouts analysis never promotesTraffic and pods look healthy, yet automated promotion never happens because the analysis provider keeps evaluating metrics for an old release label set.CI/CDAdvanced22 minProCICD-061Argo CD app-of-apps sync waves leave the platform upgrade stuck on a child dependencyThe root application looks healthy, but one child app never becomes ready because the sync ordering assumes a dependency finished before its CRDs actually became usable.CI/CDAdvanced24 minProK8S-052Readiness probe fails only after service mesh sidecar intercepts the health portThe container answers correctly on its native port, but the probe fails because the sidecar path or redirected port does not match the probe expectation.KubernetesAdvanced24 minProK8S-035Validating webhook rejects new workloads after certificate rotationThe webhook service is reachable, but admissions fail because the CABundle in the configuration no longer matches the serving certificate chain.KubernetesAdvanced27 minProK8S-054Volume attach succeeds but the pod never mountsThe control plane reports the volume attachment as healthy, yet the pod stays stuck on a subset of nodes because the node-side plugin is not present everywhere.KubernetesAdvanced27 minProK8S-060HPA and PodDisruptionBudget combine to deadlock a low-replica rolloutAutoscaling and disruption control each look reasonable alone, but together they prevent the deployment from making forward progress during an update.KubernetesAdvanced28 minProK8S-039StatefulSet scale-down keeps PVC data that contaminates the next ordinal reuseA later scale-up reuses the ordinal and attaches an old volume, causing the application to start with stale state that does not match the new deployment intent.KubernetesAdvanced28 minProK8S-094CronJob concurrency policy skips every runThe schedule is correct, but no new work starts because the previous Job remains non-terminal and blocks the concurrency guard forever.KubernetesAdvanced16 minProK8S-025Cross-namespace DNS lookup works in debug pod but not app podA quick debug shell resolves the name, while the workload container fails because search domains and runtime image differ.KubernetesIntermediate20 minProK8S-079Namespace deletion hangsMost resources are gone, but deletion never completes because one finalizer must call an API service that is no longer reachable in the cluster.KubernetesAdvanced20 minProK8S-074Admission latency spikesThe webhook service exists, but requests from one path stall because the backing Deployment lost zonal coverage and the API server keeps timing out on remote retries.KubernetesAdvanced21 minProK8S-099API server audit logs show the patch but the controller cache never sees itThe update reaches the API, yet the controller behaves as if nothing changed because its watch stream state drifted and the resync never caught up.KubernetesAdvanced22 minProK8S-070metrics-server loses the aggregated API after the front-proxy CA bundle driftsThe pod is running, but autoscaling and kubectl top fail because the aggregated API trust chain no longer matches the API server front-proxy configuration.KubernetesAdvanced23 minProK8S-051StatefulSet recovery stallsOnly one ordinal is unhealthy, but the whole recovery path stops because PodManagementPolicy still enforces ordered readiness.KubernetesIntermediate20 minProK8S-031Slow-starting JVM pod loopsLiveness and readiness probes look reasonable for a steady-state service, but the application restarts repeatedly during warmup because it never gets a dedicated startup window.KubernetesIntermediate21 minProK8S-022Service endpoint count drops during rolling updateService endpoint count drops (Rollout Stuck) is a hands-on troubleshooting drill. Pods are healthy one by one, but the endpoint pool shrinks too far during rollout and causes short traffic gaps. Kubernetes Ingress and Traffic needs to be checked by narrowing scope, recent chan...KubernetesIntermediate23 minProK8S-1679A PodDisruptionBudget exists and one drain still blocksA node drain still blocks even though replicas look healthy after a label cleanup.KubernetesIntermediate12 minProK8S-1689A PodDisruptionBudget exists and one drain still blocksA node drain still blocks after the rollout labels were standardized.KubernetesIntermediate9 minProK8S-1699A PodDisruptionBudget exists and one drain still blocksA node drain still blocks after rollout labels were cleaned up.KubernetesIntermediate9 minProK8S-1719A PodDisruptionBudget exists and one drain still blocksA drain still blocks even though a PodDisruptionBudget exists.KubernetesIntermediate9 minPro