Cluster Networking and Service Discovery
35 incident problems about Cluster Networking and Service Discovery. Start with the reviewed ones.
Read first
When the ConfigMap changed but the Pod keeps using the old valuesAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when values don't refresh after a ConfigMap update because of envFrom, subPath, volume projection, or rollout trigger differences.Kubernetes3 min readWhen CoreDNS is Running but only DNS lookups failAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when the CoreDNS Pod looks healthy but failures occur in the kube-dns path, upstream, node-local-dns, or policy.Kubernetes3 min readWhat to check first when CrashLoopBackOff appearsAn InfraTree guide that lays out the first signals to check, the CLI verification order, common misdiagnoses, and a safe recovery path when a Pod restarts repeatedly and the real cause is left in the previous logs and events.Kubernetes3 min read
Recommended problems
Reviewed problems first, then problems with detailed scenarios.
All problems (35)
Service calls fail: the endpoints list is emptyService calls fail: the endpoints list is empty is a hands-on troubleshooting drill. When a Service has no endpoints, compare its selector with the pod labels. cluster-networking-and-service-discovery needs to be checked by narrowing scope, recent change, and the current live...ReviewedKubernetesBeginner3 minFreeK8S-114NetworkPolicy allows the namespace selector but cluster DNS still failsThe app can reach peer services, yet name resolution breaks because the policy only modeled TCP while the resolver path still needs UDP.KubernetesIntermediate15 minFreeK8S-103ExternalDNS creates the new record but readiness still failsDNS updates succeed outside the pod, yet the app never heals because its resolver library keeps using stale answers through the rollout.KubernetesIntermediate17 minFreeK8S-108Service internalTrafficPolicy Local blackholes requests after a scale down leaves one node without local endpointsThe service stays healthy overall, but clients on one node get failures because the policy forbids cross-node forwarding after endpoint placement changes.KubernetesIntermediate16 minFreeK8S-131A NetworkPolicy allows the outbound proxy but still blocks OCSP and CRL endpoints, so strict clients fail external TLS validationEgress seems mostly open, yet revocation checks cannot complete because only the primary proxy path was modeled.KubernetesAdvanced17 minProK8S-105Ingress canary header routing works for HTTP but gRPC requests ignore the split and all traffic stays on stableThe progressive delivery rule appears valid, but the gRPC path follows a different routing evaluation than the header-based HTTP test path.KubernetesAdvanced17 minProK8S-146A Gateway API route binds to the correct listener, but the backendRef points at a Service in another namespace without the required ReferenceGrantEverything looks connected until cross-namespace security rules are evaluated.KubernetesAdvanced18 minProK8S-312An externalTrafficPolicy Local design is correctAn externalTrafficPolicy Local design is correct focuses on cluster-networking-and-service-discovery and asks the reader to isolate Timeouts and Latency in AWS. 실무에서는 timeouts-and-latency 증상만 보고 Pod 하나에 매달리지 말고 이벤트, 이전 로그, Service/Endpoint, 최근 배포 변경을 한 번에 묶어 보는 편이 오진을 줄입니다.KubernetesAdvanced18 minProK8S-151The ingress controller trusts X-Forwarded-Proto from the external load balancer, but an internal hop rewrites it and secure redirects begin loopingTLS is terminated correctly, yet downstream protocol awareness is now inconsistent across hops.KubernetesAdvanced18 minProK8S-192A route binds to the right listener while a cross-namespace backend reference is still unauthorized during a failover rehearsalThe object graph looks connected and the security boundary still prevents end-to-end flow. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced19 minProK8S-393A multi-cluster failover record updates correctly while clients pinned to one HTTP/2 connection never rediscover the new backend set during a staged decommissionName resolution converges and long-lived streams keep the old world alive. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.KubernetesAdvanced17 minProK8S-137An external load balancer health check hits the NodePort before readiness gating is effective, and nodes enter rotation with no viable backendsInfrastructure sees open ports, but the application is not actually ready to serve traffic yet.KubernetesAdvanced17 minProK8S-126EndpointSlice hints keep steering traffic to a drained zone until the controller catches up, and clients see retries long after evacuationDrain activity succeeded operationally, yet topology-aware routing continues sending some traffic into the old fault domain.KubernetesAdvanced17 minProK8S-129Topology-aware routing and sticky sessions combine to create a hotspot on one zone after a scale eventNeither feature is individually wrong, but together they pin too much traffic to a shrinking endpoint subset.KubernetesAdvanced17 minProK8S-264A source-range rule trusts node addresses while externalTrafficPolicy Local leaves one zone with no eligible ingress path during a failover rehearsalThe load balancer policy is correct and locality changes which nodes can actually receive traffic. Normal traffic masked the issue until the standby or alternate path became active under rehearsal conditions.KubernetesAdvanced18 minProK8S-111Gateway API reports the HTTPRoute as accepted but the listener hostname mismatch prevents any real attachmentStatus conditions look promising, yet traffic never arrives because the route was accepted by the controller but not bound to the intended hostname listener.KubernetesAdvanced18 minProK8S-116kube-proxy in IPVS mode retains a stale destination after a rapid rollout and some clients keep hitting terminated podsThe Service endpoints update correctly, yet a subset of traffic still lands on dead backends because the node-level forwarding table lagged behind the event stream.KubernetesAdvanced18 minProK8S-134Node-local DNS cache serves NXDOMAIN for a Service that was created moments later, and one node keeps the bad answer through the rolloutThe Service exists now, but one resolver path still trusts the earlier negative cache result.KubernetesIntermediate15 minProK8S-363A gateway API route is admitted while one backend reference still points to a namespace alias only the old ingress controller understood during a staged decommissionThe route object looks valid and one controller-specific assumption no longer applies. The service still works through the primary path, but one dependency only fails when the old component is finally drained away.KubernetesIntermediate16 minProK8S-144A headless Service returns pod A records correctly, but one Java client caches the first answer forever and never balances after scale-outCluster DNS is accurate, yet one library's lookup behavior defeats the intended design.KubernetesIntermediate16 minPro