Symptom20 problems· 7 reviewed

CrashLoop and Restarts

20 incident problems that show up as “CrashLoop and Restarts”.

Read first

Recommended problems

Reviewed problems first, then problems with detailed scenarios.

K8S-001Organizing the first checkpoints for a CrashLoopBackOff PodLearn in what order to check logs, events, and config when looking at a Pod stuck in a restart loop.ReviewedKubernetesBeginner18 minFreeK8S-004ImagePullBackOff caused by a missing private registry credentialA situation where the image address is correct but a missing pull secret causes failure in only certain namespaces.ReviewedKubernetesBeginner17 minFreeK8S-011A readiness probe path typo keeps the rolling update from finishingA readiness probe path typo keeps the rolling update from finishing is a hands-on troubleshooting drill. A situation where the application is fine but only the probe path is wrong, so new Pods never become ready. Kubernetes Config and Rollouts needs to be checked by narrowing...ReviewedKubernetesBeginner16 minFreeCrashLoopBackOff: pods keep restarting after a new releaseCrashLoopBackOff: pods keep restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read the previous container's log to find why a pod keeps restarting. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current liv...ReviewedKubernetesBeginner3 minFreeExit code 137 (OOMKilled): a container restarts under loadExit code 137 (OOMKilled): a container restarts is a hands-on troubleshooting drill. Learn what OOMKilled and exit code 137 mean. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는 CrashLoop...ReviewedKubernetesBeginner3 minFreeOOMKilled: a container keeps restarting after hitting its memory limitOOMKilled: a container keeps restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read OOMKilled and exit code 137, then fix memory limits and the workload that exceeds them. kubernetes-workload-reliability needs to be checked by narrowing scope, recent cha...ReviewedKubernetesBeginner15 minFree

All problems (20)

CrashLoopBackOff: pods keep restarting after a new releaseCrashLoopBackOff: pods keep restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read the previous container's log to find why a pod keeps restarting. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current liv...ReviewedKubernetesBeginner3 minFreeExit code 137 (OOMKilled): a container restarts under loadExit code 137 (OOMKilled): a container restarts is a hands-on troubleshooting drill. Learn what OOMKilled and exit code 137 mean. kubernetes-workload-reliability needs to be checked by narrowing scope, recent change, and the current live signal before rollback. 실무에서는 CrashLoop...ReviewedKubernetesBeginner3 minFreeLINUX-005A missing WorkingDirectory problem that puts a systemd service into a restart loopA situation where the application itself is fine but a wrong unit-file path setting kills the service.ReviewedLinuxIntermediate20 minFreeK8S-011A readiness probe path typo keeps the rolling update from finishingA readiness probe path typo keeps the rolling update from finishing is a hands-on troubleshooting drill. A situation where the application is fine but only the probe path is wrong, so new Pods never become ready. Kubernetes Config and Rollouts needs to be checked by narrowing...ReviewedKubernetesBeginner16 minFreeK8S-004ImagePullBackOff caused by a missing private registry credentialA situation where the image address is correct but a missing pull secret causes failure in only certain namespaces.ReviewedKubernetesBeginner17 minFreeK8S-001Organizing the first checkpoints for a CrashLoopBackOff PodLearn in what order to check logs, events, and config when looking at a Pod stuck in a restart loop.ReviewedKubernetesBeginner18 minFreeOOMKilled: a container keeps restarting after hitting its memory limitOOMKilled: a container keeps restarting (CrashLoop and Restarts) is a hands-on troubleshooting drill. Read OOMKilled and exit code 137, then fix memory limits and the workload that exceeds them. kubernetes-workload-reliability needs to be checked by narrowing scope, recent cha...ReviewedKubernetesBeginner15 minFreeK8S-024Pod restartsThe file exists and the volume is mounted, but the runtime user cannot read the config because the mode is too strict.KubernetesBeginner15 minFreeLINUX-036Journald rate limiting hides a flapping service patternThe service restarts repeatedly, but the most useful logs never appear because journald is suppressing the repeated messages during the worst part of the incident.LinuxIntermediate20 minFreeLINUX-020The env-var file was edited, but an already-running process does not know the valueThe env-var file was edited, but an already-running process does not know the... is a hands-on troubleshooting drill. The most basic operational problem for understanding the difference between a config file change and a process restart. Linux Service Operations needs to be ch...LinuxBeginner14 minFreeK8S-002Analyzing a service failure where the Pod is Running but receives no trafficA scenario for step-by-step checking of Service, Endpoint, readiness probe, and selector mismatch possibilities.KubernetesIntermediate22 minProK8S-014An overly aggressive liveness probe keeps restarting even healthy PodsAn overly aggressive liveness probe keeps restarting even healthy Pods is a hands-on troubleshooting drill. A situation where the liveness criteria become too tight during periods of high application load. Kubernetes Config and Rollouts needs to be checked by narrowing scope,...KubernetesIntermediate19 minProK8S-003Interpreting backoffLimit in a batch workload where the Job never finishesAn advanced problem that connects the batch failure retry policy with the container exit code to find the root cause.KubernetesAdvanced27 minProSECURITY-006Audit log volume drops after agent restart despite healthy daemon statusThe collector service looks healthy, but filtering or delivery state changes silently reduce security log coverage.SecurityAdvanced27 minProK8S-052Readiness probe fails only after service mesh sidecar intercepts the health portThe container answers correctly on its native port, but the probe fails because the sidecar path or redirected port does not match the probe expectation.KubernetesAdvanced24 minProK8S-039StatefulSet scale-down keeps PVC data that contaminates the next ordinal reuseA later scale-up reuses the ordinal and attaches an old volume, causing the application to start with stale state that does not match the new deployment intent.KubernetesAdvanced28 minProK8S-029CSI driver recovers volume attach only after controller restartPersistent volume attachment gets stuck until the controller is restarted, pointing to an unhealthy reconciliation loop.KubernetesAdvanced29 minProLINUX-022Systemd service restarts too quickly to capture useful logsA restart loop makes the service hard to inspect because the process dies before operators can capture the right signal.LinuxIntermediate21 minProK8S-031Slow-starting JVM pod loopsLiveness and readiness probes look reasonable for a steady-state service, but the application restarts repeatedly during warmup because it never gets a dedicated startup window.KubernetesIntermediate21 minProK8S-040Node eviction startsApplication storage looks healthy on the persistent volume, but a temporary working directory on the node root filesystem silently fills and triggers eviction pressure.KubernetesIntermediate24 minPro