Fast failover becomes unstable after a QoS change even though throughput tests show the path is otherwise healthy.
A BFD-assisted routing failover flaps repeatedly because the underlay policing rate treats bursty control traffic as expendable even while data traffic remains healthy
Applications mostly work, yet the control plane oscillates because microbursts of BFD or routing packets exceed a policy that was tuned only with data-plane testing.
Scenario
What to check first
- Identify the primary failure signal in the Healthy Data Plane, Policed Control Burst scenario.
- Separate visible symptoms from the underlying technical dependency.
- Describe the safest recovery path and the follow-up prevention work.
Checking checklist
- Summarize the current impact and the last known change.
- Collect direct evidence from logs, runtime state, and configuration before changing anything.
- Separate immediate recovery from permanent prevention work.
Recovery and prevention
Validate control-plane packet treatment under burst conditions before raising protocol timers.
Questions worth viewing together
Community-field network problem inspired by routing and QoS threads where BFD or routing control traffic was policed away even while the data plane look... Control traffic often fails under a different burst and latency profile than ordinary data traffic.
Teams often blame an unstable peer when local policing is selectively harming the control protocol.
QoS tuned around throughput alone can quietly break fast-routing protocols that depend on bursty control traffic.
Similar cases seen in the field