At 23:19, two iBGP links from the Frankfurt router failed within the same minute. Both ran over WireGuard, both remained configured and administratively up, and the other routing sessions on the router stayed established.
The wallboard made the shape of the failure obvious. Four directed BGP sessions were down across two bilateral links, while the routers, transit sessions and unrelated overlay paths remained healthy.

The incident window. “Four sessions down” is the two failed adjacencies viewed from both ends.
The traffic graphs ruled out congestion almost immediately. In the 24 hours before the failure, the links were carrying little more than routing keepalives:
| Link, measured at Frankfurt | Average receive | Average transmit | Peak receive | Peak transmit |
|---|---|---|---|---|
| toward Manchester | 52 bit/s | 172 bit/s | 490 bit/s | 754 bit/s |
| toward Zurich | 41 bit/s | 42 bit/s | 566 bit/s | 446 bit/s |
The final hour before the incident looked the same. There was no rising load, error burst or queue pressure before the sessions disappeared. Looking back over seven days did find some short UDP benchmark runs: the larger link briefly reached about 0.77 Mbit/s and the smaller one about 57 kbit/s. Those tests had finished well before this incident and were nowhere near enough traffic to explain it.
The useful evidence was in the five-minute WireGuard byte deltas after the failure:
| Tunnel endpoint | Received | Transmitted |
|---|---|---|
| Frankfurt side toward Manchester | 0 bytes | about 8.2 kB |
| Manchester side toward Frankfurt | 0 bytes | about 8.6 kB |
| Frankfurt side toward Zurich | about 7.9 kB | about 5.2 kB |
| Zurich side toward Frankfurt | 0 bytes | about 8.2 kB |
Those oddly specific small numbers were the best clue in the whole incident. A WireGuard handshake initiation is 148 bytes and a response is 92 bytes. Retried every five seconds, they produce roughly the byte rates in the graph once Prometheus scrape interpolation is allowed for.
On the Manchester link, both ends were sending initiations and neither end received anything. On the Zurich link, initiations reached Frankfurt and Frankfurt sent the smaller responses, but none of those responses reached Zurich. Packet captures agreed with the counters.
That put the failure beyond the local tunnel configuration. It looked like selective underlay loss: complete in one case and return-path-only in the other. I cannot prove what an upstream network did from these measurements alone, but I could rule out a bad BGP policy, a full link, a router reboot and a general WireGuard failure.
The BGP graphs lagged the packet failure by their timers. The first endpoint reported down at 23:20, and the final affected session followed at 23:22. Prometheus, the routing exporters and every unaffected peer remained healthy throughout. A transmit-drop counter on one dead tunnel started climbing several minutes later, which was a consequence of packets accumulating behind a handshake that never completed rather than the cause.
Then, at about 23:47, the missing packets started arriving again. All four BGP sessions re-established by 23:48 without a configuration change. The outage lasted roughly twenty-nine minutes.
An interface being “up” was almost useless here. The combination of BGP state, per-direction byte counters, packet sizes and a wallboard which kept unrelated paths visible told the story much more precisely: two control-plane links failed together, but they did not fail in the same direction.