A scenario that narrows down the root cause, centered on designing a retry policy that distinguishes transient errors from structural ones, in the situation of a deployment pipeline that only piles up cost with endless retries on a transient error.
A deployment pipeline that only piles up cost with endless retries on a transient error
A problem about judging which errors should actually be retried in a pipeline where retry settings have grown indiscriminately.
시나리오
단서
구독하면 이어서 볼 수 있어요
이 문제의 전체 시나리오와 점검 체크리스트, 복구 순서, 모범 풀이는 Pro 구독에서 열립니다.
먼저 볼 것
- Identify the primary failure signal in the Pipeline Policy scenario.
- Separate visible symptoms from the underlying technical dependency.
- Describe the safest recovery path and the follow-up prevention work.
점검 체크리스트
- Summarize the current impact and the last known change.
- Collect direct evidence from logs, runtime state, and configuration before changing anything.
- Separate immediate recovery from permanent prevention work.
복구와 재발 방지
Choose the smallest safe recovery action first, then record the prevention work that reduces repeat incidents.
같이 보면 좋은 질문
Blind retries can rerun destructive steps, consume runner capacity, and bury the original transient-vs-persistent failure signal.
Decide whether the failure is safe to retry, then isolate the exact step and side effect before re-execution.
현장에서 본 비슷한 케이스