Typical symptom
When Preview passes but the secret is missing only under the production feature flag is rarely a single wrong setting — it usually appears when the deploy boundary, runtime state, cache, permissions, and network path drift together. Users arriving from search should first narrow the symptom, using this guide's search intent — Targets search intents like preview passed but production secret missing, feature flag exposed missing secret, and deploy success but prod only secret path empty. — as the basis for the first signals to check.
Early on, rather than attempting a full rollback or a blind restart, check which node, Pod, job, user, or path the blast radius is tied to. If the scope is narrow, compare recent changes against the healthy state; if it is wide, start from the shared dependencies.
Signals to check first
The first thing to look at is not the last error line but the boundary where the same failure recurs. Grouping the problem around signals like preview passed prod failed, feature flag missing secret, secret path empty only in prod, runtime config mismatch narrows the root-cause candidates even when the logs are long.
- First check whether preview and production look at the same secret source.
- Distinguish whether a feature flag or release toggle opened a different code path.
- Don't treat runtime secret injection and build-time config substitution as the same thing.
- Even if the secret name is the same, compare whether the namespace, account, or region differs.
- View side by side the config snapshots referenced by a healthy preview request and a failing production request.
Logs and CLI examples
The commands below don't hand you the answer directly; they are the first observation points for narrowing the cause. Comparing their output against a known-good point in time or a healthy resource with the same role cuts down time spent just retrying.
gh run view <run-id> --log
git diff --name-only <previous>..<current>
Common misdiagnoses
The most dangerous pattern in operational incidents is mistaking the symptom name for the cause. The same timeout, permission denied, or rollout failure can have its real cause in a different layer — cache, permission inheritance, Secret scope, stale client connections, or proxy headers.
- Assuming that because preview passed, the secret takes the same path.
- Redeploying only the secret store itself without looking at the code path the flag opens.
Safe recovery order
Recovery starts at the smallest unit. First pin the current state with read-only checks, then verify changes on a limited-impact resource. Hard-to-reverse actions like a full service restart, clearing the entire cache, or relaxing security policy should be chosen only after the root-cause candidates are narrowed.
- This type comes from environment-contract differences more often than application bugs, so the deploy team and the app team may be looking at different evidence.
- When the preview path is healthy but only the production flag path differs, the logs look sparse and the initial response is delayed.
Prevention
After an incident ends, record "why that state lingered" rather than just a one-line cause. Check whether there was a gap between automation and operational procedure — in the deploy pipeline, runtime reload, permission inheritance, certificate renewal, or network policy.
Following the related hubs and problems lets you re-diagnose the same symptom in other environments.
Related InfraTree problems
The problems below are public exercises for practicing this guide as real scenarios. Solving them after reading lets you practice splitting signals first and writing out the recovery direction as sentences.