← 문제 라이브러리
CI/CD Intermediate CICD-004 · 18 min

A deployment pipeline that only piles up cost with endless retries on a transient error

A problem about judging which errors should actually be retried in a pipeline where retry settings have grown indiscriminately.

무료
시나리오

A scenario that narrows down the root cause, centered on designing a retry policy that distinguishes transient errors from structural ones, in the situation of a deployment pipeline that only piles up cost with endless retries on a transient error.

먼저 볼 것
  • Identify the primary failure signal in the Pipeline Policy scenario.
  • Separate visible symptoms from the underlying technical dependency.
  • Describe the safest recovery path and the follow-up prevention work.
점검 체크리스트
  1. Summarize the current impact and the last known change.
  2. Collect direct evidence from logs, runtime state, and configuration before changing anything.
  3. Separate immediate recovery from permanent prevention work.
복구와 재발 방지

Choose the smallest safe recovery action first, then record the prevention work that reduces repeat incidents.

같이 보면 좋은 질문
Why can pipeline retries make an incident worse?

Blind retries can rerun destructive steps, consume runner capacity, and bury the original transient-vs-persistent failure signal.

What is the right first decision?

Decide whether the failure is safe to retry, then isolate the exact step and side effect before re-execution.