Topic65 problems· 4 reviewed

Linux Service Operations

65 incident problems about Linux Service Operations. Start with the reviewed ones.

All problems (65)

LINUX-093Systemd path unit keeps relaunching a failed helperThe service appears to restart itself, but the real trigger is a path unit that keeps firing on the same file activity loop.LinuxAdvanced17 minProLINUX-090Systemd cgroup memory limit kills the helper process and the main service logs look unrelatedThe application seems to fail in an unrelated code path, but the hidden cause is a helper process being OOM-killed inside the unit's cgroup budget.LinuxAdvanced18 minProLINUX-078nftables set refresh leaves stale IP entriesThe firewall update appears to run, but some stale members stay active because the atomic update path failed halfway and the old set was never swapped cleanly.LinuxAdvanced19 minProLINUX-087Package rollback restores binaries but not the migrated systemd unit dependency graphThe older package version is back, yet the service still behaves incorrectly because the dependency units changed in a later release and were not reverted with the package.LinuxAdvanced20 minProLINUX-067chrony source priority drift pushes one site out of acceptable Kerberos time skewNTP is technically running everywhere, but one location follows a lower-quality source and drifts just far enough to break time-sensitive authentication.LinuxAdvanced21 minProLINUX-054nftables hotfix disappearsA manual packet filter change solves the immediate outage, but the next reload or reboot brings back the old policy because the permanent source of truth was never updated.LinuxAdvanced23 minProLINUX-007TLS verification intermittently failing due to system time driftCovers a system-time problem where the application is fine but certificate validity-time checks drift.LinuxAdvanced23 minProLINUX-035PAM limits are raised but the systemd service still hits too many open filesInteractive shells inherit the new limit correctly, but the daemon still fails under load because its service-level limit remains unchanged.LinuxAdvanced25 minProLINUX-026Swap activity stays high after memory spike is goneSwap activity stays high (Timeouts and Latency) is a hands-on troubleshooting drill. The incident is over, but swap churn remains elevated and keeps latency unpredictable for the rest of the day. Linux Service Operations needs to be checked by narrowing scope, recent change, a...LinuxAdvanced25 minProLINUX-023Ephemeral disk pressure kills cache but not the root causeEphemeral disk pressure kills cache but not the root cause is a hands-on troubleshooting drill. Cleaning temp files buys time, but write amplification from another process quickly recreates the same disk pressure. Linux Service Operations needs to be checked by narrowing scope...LinuxAdvanced26 minProLINUX-1489A PAM faillock reset succeeds and remote users still failUsers remain locked out only on nodes with home-mounted identity caches after a faillock reset.LinuxAdvanced10 minProLINUX-1482A Fedora rsyslog queue drains and journald still growsDisk use keeps growing under journald after rsyslog spool hardening on Fedora-based log collectors.LinuxAdvanced11 minProLINUX-1486A sudoers policy looks correct and one host still denies adminsAdmin sudo fails only after package upgrades on hosts bootstrapped with cloud-init snippets.LinuxAdvanced11 minProLINUX-1487An mdadm reshape resumes after reboot and write performance collapsesThroughput craters after a reboot in the middle of an array reshape even though the array reports healthy.LinuxAdvanced12 minProLINUX-1484An NFS mount remounts cleanly and one app still sees stale handlesA remounted NFS path still throws stale file handle errors only for processes that survived the maintenance window.LinuxAdvanced12 minProLINUX-092Kernel parameter is tuned in sysctl.d but an earlier file winsThe desired setting exists on disk, yet the host still uses another value because an earlier or later file in the sysctl load order overrides it.LinuxAdvanced15 minProLINUX-028Container host shows low CPU while one cgroup throttles hardContainer host shows low CPU (Resource Exhaustion) is a hands-on troubleshooting drill. Overall node utilization looks safe, but one service is heavily throttled because its cgroup limit is too low. Linux Service Operations needs to be checked by narrowing scope, recent change...LinuxIntermediate20 minProLINUX-022Systemd service restarts too quickly to capture useful logsA restart loop makes the service hard to inspect because the process dies before operators can capture the right signal.LinuxIntermediate21 minProLINUX-025Socket backlog tuning hides application accept bottleneckKernel queue settings were increased, but the service still stalls because the application cannot drain connections fast enough.LinuxIntermediate22 minProLINUX-003Tracking the upstream application bottleneck in an Nginx 502 incidentA troubleshooting exercise that comprehensively considers the proxy, ports, application state, and timeout settings.LinuxAdvanced28 minPro